Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
85
datasets available to search
ShareScore release 0.9.0
Dataset results
85 results for “lexicon”
Italian Verb Lexicon for Sentiment Inference
<p><strong>Italian Verb Lexicon for Sentiment Inference</strong></p> <p><strong>Theory:</strong></p> <p>For a description of the theory behind the specifications of the corpus, please read the attached paper. </p> <p><br> <strong>Example of json entry:</strong></p> <p>{"verb": "soddisfare", "frames": [{"fillers": ["Subj", "DirObj/IndObj"], "polarity": "POS", "effects": [["DirObj/IndObj", "pos"]], "expectations": [], "examples": ["L'offerta ha soddisfatto i clienti.", "Soddisfare al pubblico."], "remarks": [], "relations": [["Subj", "DirObj/IndObj", "pro"]]}]}</p> <p><strong>Description: </strong><br> The verb "soddisfare" has 2 frames, a subject followed by a direct or indirect object. The verb polarity is positive. There is an positive effect on the direct (indirect) object. No expectations. There is a in favour (pro) relation from the subject to the direct (indirect) object. Two example sentenes are given.</p> <p><strong>Synonyms:</strong><br> Some entries are references to synonym verbs with identical frames:</p> <p>{"verb": "consacrare", "germanTranslation": "widmen", "frameReference": "dedicare", "examples": ["Consacrare tempo alle sue passioni"]}</p> <p>Here, "cosacrare" and "dedicare" are assumed synonyms with the same syntactic frames.</p> <p><strong>Used Tags:</strong></p> <p>A few explanations on the tags used in the verb specifications sheets.</p> <p>Subj = subject</p> <p>DirObj = direct object, as in "Il professore legge __il giornale__".</p> <p>IndObj = indirect object, as in "Permettere qualcosa __a qualcuno__".</p> <p>RefObj = reflexive object (pronoun), as in "La squadra avversaria __si__ è arrabbiata moltissimo". </p> <p>PrepObj[prep] = prepositional phrase; the preposition is specified in the square brackets. If more than one preposition can occur,<br> no specification is given.</p> <p>SubCl = a subordinate clause, usually introduced by "che" or "di" such as in "Ha detto __di andarsene__", <br> "Ha detto __che tutto è andato bene__".</p> <p>mod = any type of modifier, mostly adverbs, e.g. "Se ne è andato __subito__".</p> <p><br> *, e.g. mod* = indicates optionality</p> <p><br> </p> <p> </p>
Bulu Puroik Lexicon
<p>The Bulu Puroik lexicon contains:</p> <p>- 5081 plain text files with orgmode markup (./lemmas/org/).</p> <p>- 5879 audio files (./sounds/)</p> <p>- 306 image files (./images/)</p> <p>The org-file names correspond to the identifier given in the dictionary published in the appendix of "A Grammar of Bulu Puroik" (<a href="https://boristheses.unibe.ch/2251/">https://boristheses.unibe.ch/2251/</a>).</p> <p>The lexicon can be searched with rgrep. The sound files can be played with org-player.el</p> <p>This work is licensed under a <a href="https://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a></p> <p> </p>
VnEmoLex: A Vietnamese emotion lexicon for sentiment intensity analysis
<p>VnEmoLex is the moderate-sized data set annotated for eight basic emotions: joy, sadness, anger, fear, trust, disgust, surprise and anticipation for Vietnamese. It is built on the NRC Word-Emotion Association Lexicon (EmoLex)<sup>1 </sup>and the Viet Wordnet<sup>2</sup> . VnEmoLex has total 12,795 words of which 4431 words are from the EmoLex dictionary, 8364 words are taken from the Viet Wordnet.</p> <p><sup>1 </sup>http://saifmohammad.com/WebPages/NRC-Emotion-Lexicon.htm</p> <p><sup>2 </sup>http://http://viet.wordnet.vn/wnms/</p>
VeLeCa: a Verbal Lexicon of Catalan
<p>VeLeCa is an inflected lexicon of Catalan verbal inflection containing the phonological form of 174,200 word forms from 3,484 lexemes and their respective lexical and morphosyntactic values and frequencies.</p>
VeLeRo: a Verbal Lexicon of Romanian
<p>VeLeRo, an inflected lexicon of Standard Romanian which contains the full paradigm of 7297 verbs in phonological form, with lemma-level and cell-level frequencies.</p> <p>Full info at: Herce, B., Pricop, B. VeLeRo: an inflected verbal lexicon of standard Romanian and a quantitative analysis of morphological predictability. <em>Lang Resources & Evaluation</em> (2024). https://doi.org/10.1007/s10579-024-09721-3</p>
Lubrang Brokpa Lexicon - basic word list
<p>These sound files constitute the elicitation of the lexical entries of the Basic Word List in the Lubrang variety of Brokpa. Lubrang village is a recent (early 20th century) settlement of Brokpa speaker originating in Sakteng village of Bhutan. They left Bhutan due to its heavy taxation of the semi-nomadic Brokpa households and its anti-Gelukpa policies and settled in the then Tibetan-administered area on land belonging to the Khispi people of Lish village, to which they continue to pay an annual tax. Lubrang Brokpa should hence be close to Merak and Sakteng (Bhutan) Brokpa, and not so close to Nyukmadung and Senge Brokpa spoken closer by. Because of the speaker’s paternal background there may be some admixture with Dirang Tshangla.</p>
Dataset: Lexicon Pharmaceuticals, Inc. (LXRX) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Sherdukpen Lexicon - Shergaon
<p>These files contain the original sound files, the cut sound files, and the transcriptions of the lexicon of Sherdukpen collected from Shergaon.</p> <p>Bodt, Timotheus Adrianus. 2024. <em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><span><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&code=view&bookID=146">LANGUAGE AND LINGUISTICS > (sinica.edu.tw)</a></span></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>
Sherdukpen Lexicon - Rupa
<p>This dataset contains all the sound files and transcribes files of the Rupa variety of Sherdukpen. </p> <p>Bodt, Timotheus Adrianus. 2024. <em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><span><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&code=view&bookID=146">LANGUAGE AND LINGUISTICS > (sinica.edu.tw)</a></span></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>
Duhumbi Lexicon
<p>This supplementary material discusses and presents issues related to the Duhumbi lexicon. There are sound files, descriptions, and analysis. There are two descriptive files, one discussing the inherited nouns, and one discussing the borrowed part of the Duhumbi lexicon. Then, there is a zip file with all the cut lexical sound files. There are 11 short descriptions of minimal pairs and the corresponding sound files in zip files. Then, there are five zip files with sound recordings. These files are described in the tables in the overview file.</p> <p>Bodt, Timotheus Adrianus. 2024. <em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><span><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&code=view&bookID=146">LANGUAGE AND LINGUISTICS > (sinica.edu.tw)</a></span></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: bodttim (at) gmail (dot) com</p>
A Kam Lexicon
<p>This dataset contains a lexicon of Kam [ISO 639-3: kdx; Glottocode: kamm1249], a hitherto difficult to classify Niger-Congo language spoken in Central-Eastern Nigeria (see Lesage 2018a,b, 2019a,b,c, in preparation for more on the language). It is mainly meant as a source to cite for RefLex (Segerer & Flavier 2011-2018), where the word list is also available. Only word forms, glosses, and word class information are available for now. There is no introduction so far. I intend to publish a more comprehensive lexicon or dictionary of Kam, with introduction about the phonology and grammar, in the future. Note that, contrary to general Africanist tradition, <j> represents a palatal approximant, <dʒ> representing voiced palatoalveolar affricates, and <tʃ> represents voiceless palatoalveolar affricates instead of <c>.</p>
Pere lexicon
<p>A lexicon of Pere (Bere, Mbre) of Côte d’Ivoire. Pere (in the literature also spelled Pɛrɛ, Bere, Mbre) is a seriously endangered language of central Côte d’Ivoire. It is listed as “Mbre” in Glottolog (viewed July 2018), code mbre1244, and in ISO 639-3, code mka.</p>
Inflected lexicon of Russian Nouns in IPA notation
<p>This inflected lexicon of Russian Nouns is based on data generated by a DATR fragment for the nominal system of Russian (Dunstan Brown et al, 2011), which was used in Brown and Hippisley (2012). The files were automatically transcribed to an IPA-based notation by Sacha Beniamine (2018), following rules provided by Dunstan Brown. We also corrected a few errors in the original data found manually. We thank Sasha krasovitsky for providing comments on the phonological feature definitions.</p>
CLDF dataset derived from Luangthongkum's "Proto-Karen Phonology and Lexicon" from 2019
<p>Cite the source of the dataset as:</p> <blockquote> <p>Luangthongkum, Theraphan (2019). A View on Proto-Karen Phonology and Lexicon, Journal of Southest Asian Linguistics Society, 12.1, i-lii. doi: http://hdl.handle.net/10524/52441</p> </blockquote>
AraVeLex: Modern Standard Arabic Verbal lexicon
<p>Paralex lexicon of Modern Standard Arabic Verbs, providing inflected forms in orthography and phonemic transcription.</p>
Vietic 116 item phylogenetic lexicon
<p>The file is a116 item lexicostatistical dataset for classification of the Vietic languages. The set includes 30 Vietic doculects, Proto-Vietic, plus Khmu and Jahai as out-groups. Included is a listing of the sources, and the NEXUS file with our cognate value assignments, which we created to run on SplitsTree to generate phylograms and NeighborNets. The 116-item list was the outcome of beginning with the Swadesh 100 and 200 lists and reconciling these with the available data with the aim of achieving at least 80% coverage for each lect in the analysis. Procedurally, sources were selected and lexicons aggregated in a spreadsheet, with rows identified with Swadesh 100 and 200 items, subject to semantic and phonological adjustments as we judged necessary. For most of the languages, full coverage of the Swadesh 100 categories was not possible, with 20 or more gaps being common. Some 40 additional categories were added from the Swadesh 200 list, based on the 40 best represented items in the aggregated data, seeking to achieve a 120-item list with at least 100 items coverage for all lects, ultimately settling on 116 items.<strong> </strong></p>
Lexicon of Polarity Shifting Directions
<p>This dataset provides information on the shifting direction of polarity shifters. Shifting directions specify whether a polarity shifter can affect <strong>only positive</strong> polar expressions, <strong>only negative</strong> ones or shift in <strong>both</strong> directions.</p> <p>We cover all shifters found in the <a href="https://doi.org/10.5281/zenodo.3545947">shifter lexicon</a> of <a href="https://doi.org/10.1017/S135132492000039X">Schulder, Wiegand and Ruppenhofer (JNLE 2021)</a>, which contains verbs, noun and adjectives.</p> <p><strong>Data</strong><br> A list of 2521 polarity shifters, labeled for their shifting direction. Contains 863 shifters that affect only positive polar expressions, 288 shifters that affect only negative polar expressions and 1370 shifters that can shift in both directions.</p> <ul> <li>File: <code>shifting_directions.txt</code></li> <li>The lexicon is a comma-separated value (CSV) table</li> <li>Each line follows the format <code>POS,LEMMA,DIRECTION_LABEL,SOURCE</code>. <ul> <li><code>POS</code>: The part of speech of the word (<code>verb</code>, <code>noun</code>, <code>adj</code>)</li> <li><code>LEMMA</code>: The lemma representation of the word in question. Multiword expressions are separated by an underscore (<code>WORD_WORD</code>).</li> <li><code>DIRECTION_LABEL</code>: Whether the shifter affects only positive polarities (<code>AFFECTS_POSITIVE</code>), only negative polarities (<code>AFFECTS_NEGATIVE</code>) or can shift in both directions (<code>AFFECTS_BOTH</code>).</li> <li><code>SOURCE</code>: Whether the word was part of the gold standard (<code>GOLD_STANDARD</code>) or was labeled automatically (<code>AUTOMATIC</code>).</li> </ul> </li> </ul> <p><strong>Attribution</strong><br> This dataset was created as part of the following publication:</p> <p><a href="http://marc.schulder.info/">Schulder, Marc</a> and <a href="http://www.coli.uni-saarland.de/~miwieg/">Wiegand, Michael</a> and <a href="http://ruppenhofer.de/">Ruppenhofer, Josef</a> (2020). <strong>"Enhancing a Lexicon of Polarity Shifters through the Supervised Classification of Shifting Directions"</strong>. <em>Proceedings of the 12th Conference on Language Resources and Evaluation (LREC)</em>, pages 5010–5016, Marseille, France, May 11-16, 2020.</p> <p>If you use the data in your research or work, please cite the publication.</p>
Bootstrapped Lexicon of English Polarity Shifters
<p>We provide a bootstrapped lexicon of English polarity shifters and their shifting direction. We cover verbs, nouns and adjectives. Our lexicon provides 2521 shifters among a vocabulary of 9145 words, taken from WordNet v3.1 (Miller et al., 1990).</p> <p>We also provide a dataset of 2631 verb phrases that are annotated for shifting polarities. The phrases are taken from the Amazon Product Review Data corpus (Jindal & Liu, 2008).</p> <p><strong>Data</strong></p> <p><strong>1. Polarity Shifter Lexicon</strong></p> <p>A list of 9145 words, annotated for whether they are polarity shifters. Contains 2631 shifters and 6514 non-shifters.</p> <ul> <li>File: <code>shifters.txt</code></li> <li>The lexicon is a comma-separated value (CSV) table</li> <li>Each line follows the format <code>POS,LEMMA,SHIFTER_LABEL,SOURCE</code>. <ul> <li><code>POS</code>: The part of speech of the word (<code>verb</code>, <code>noun</code>, <code>adj</code>)</li> <li><code>LEMMA</code>: The lemma representation of the word in question. Multiword expressions are separated by an underscore (<code>WORD_WORD</code>).</li> <li><code>SHIFTER_LABEL</code>: Whether the word is a polarity shifter (<code>SHIFTER</code>) or a non-shifter (<code>NONSHIFTER</code>)</li> <li><code>SOURCE</code>: Whether the word was part of the gold standard (<code>GOLD_STANDARD</code>) or was bootstrapped (<code>BOOTSTRAPPED</code>). All labels, both from gold standard and bootstrap output, were verified by a human annotator.</li> </ul> </li> </ul> <p><strong>2. Sentiment Verb Phrases</strong></p> <p>A set of verb phrases, annotated for the polarity of the verb phrase and the polarity of a polar noun that it contains. Can be used to evaluate whether a polarity classifier correctly recognizes polarity shifting. The file starts with 400 phrases containing shifter verbs, followed by 2231 phrases containing non-shifter verbs.</p> <ul> <li>File: <code>sentiment_phrases.txt</code></li> <li>Every item consists of: <ul> <li>The sentence from which the VP and the polar noun were extracted.</li> <li>The VP, polar noun and the verb heading the VP.</li> <li>Constituency parse for the VP.</li> <li>Gold labels for VP and polar noun by a human annotator.</li> <li>Predicted labels for VP and polar noun by RNTN tagger (Socher et al., 2013) and <code>LEX_gold</code> approach.</li> <li>Items are separated by a line of asterisks (*)</li> </ul> </li> </ul> <p><strong>Attribution</strong><br> This dataset was created as part of the following publication:</p> <p><a href="http://marc.schulder.info/">Schulder, Marc</a> and <a href="http://www.coli.uni-saarland.de/~miwieg/">Wiegand, Michael</a> and <a href="http://ruppenhofer.de/">Ruppenhofer, Josef</a> (2020). <strong>"Automatic Generation of Lexica for Sentiment Polarity Shifters"</strong>. In: <em>Natural Language Engineering</em>. <a href="https://doi.org/10.1017/S135132492000039X">doi:10.1017/S135132492000039X</a></p> <p>If you use the data in your research or work, please cite the publication.</p>
A part-of-speech (POS) lexicon of Classical Tibetan for NLP
<p>This part-of-speech (POS) lexicon of Classical Tibetan was prepared in the course of the research project 'Tibetan in Digital Communication' (2012-2015) hosted at SOAS, University of London and funded by the UK's Arts and Humanities Research Council (grant code: AH/J00152X/1). The data for verbs comes from a digitized version of <em>A Lexicon of Tibetan Verb Stems as Reported by the Grammatical Tradition</em> (Munich: Bayerische Akademie der Wissenschaften, 2010) by Nathan W. Hill. Otherwise data comes from the manually part-of-speech tagged training data produced by the corpus and a few lexical items specifically added by hand to improve rule based tagging.</p>
Phlorest phylogeny derived from Robinson and Holton 2012 'Internal Classification of the Alor-Pantar Language Family Using Computational Methods Applied to the Lexicon'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Robinson, L. C., & Holton, G. (2012). Internal Classification of the Alor-Pantar Language Family Using Computational Methods Applied to the Lexicon. Language Dynamics and Change, 2(2), 123-149. doi:10.1163/22105832-20120201</p> </blockquote>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.