Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
85
datasets available to search
ShareScore release 0.9.0
Dataset results
85 results for “Lexicon”
Finnish War Letter Emotion Lexicon
<p>The XLSX file contains (1) a manually constructed list of Finnish emotion words in the context of World War 2, (2) the frequency of each emotion word in the Finnish War Letter Collection, and (3) information on whether the emotion word is included in the filtered FEIL emotion lexicon (emotion intensity score >0.6). The construction of the Finnish War Letter Lexicon is described in: Turunen, R., Taskinen, I., Uusitalo, L. & Kivimäki, V. (2022). “Mining Emotions from the Finnish War Letter Collection, 1939–1944”. <em>Proceeding of the 6th Digital Humanities in the Nordic and Baltic Countries Conference (DHNB 2022)</em>, Uppsala, Sweden, March 15–18, 2022.</p>
Verbal aspectual lexicon for BCS: Aspect.BCS
<p>The dataset is novel language resource for retrieving and researching verbal aspectual pairs in BCS (Bosnian, Croatian, and Serbian) created using Linguistic Linked Open Data (LLOD) principles. As there is no resource to help learners of Bosnian, Croatian, and Serbian as foreign languages to recognize the aspect of a verb or its pairs, we have created a new resource that will provide users with information about the aspect, as well as the link to a verb’s aspectual counterparts. This resource also contains external links to monolingual dictionaries, Wordnet, and BabelNet. See also https://github.com/max-ionov/aspect-db/tree/main/rdf</p>
Identity Lexicon and Keyword Counts for "The Life of a Tie: Social Origins of Network Diversity"
<p>Prototype identity lexicon for the categories of occupation (e.g., "reporter" at Boston Globe), familial roles (e.g., proud "father"), political affiliation (life-long "democrat"), and cultural and sports interests (e.g., "hiphop", "NFL").</p> <p>Also includes a CSV file with counts of each identity keywords matched for all 572K users in our dataset. </p> <p> </p>
CLDF dataset derived from Sidwell and Alves' "Vietic Lexicon" from 2021
<p>Cite the source of the dataset as:</p> <blockquote> <p>Sidwell, Paul, & Alves, Mark. (2021). Vietic 116 item phylogenetic lexicon (First version (26 Aug 2021)eng) [Data set]. 9th International Conference on Austroasiatic Linguistics (ICAAL 9), Lund, Sweden. Zenodo. https://doi.org/10.5281/zenodo.5263195.</p> </blockquote>
Sartang Lexicon - Rahung
<p>This dataset contains the sound files (original and cut) and transcribed sound files of Sartang - Rahung variety.</p> <p>Bodt, Timotheus Adrianus. 2024. <em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><span><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&code=view&bookID=146">LANGUAGE AND LINGUISTICS > (sinica.edu.tw)</a></span></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>
Sartang Lexicon - Khoitam
<p>This dataset contains the sound files (original and cut) and transcribed sound files of Sartang - Khoitam variety.</p> <p>Bodt, Timotheus Adrianus. 2024. <em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><span><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&code=view&bookID=146">LANGUAGE AND LINGUISTICS > (sinica.edu.tw)</a></span></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>
Sartang Lexicon - Jerigaon
<p>This dataset contains the sound files (original and cut) and transcribed sound files of Sartang - Jerigaon variety.</p> <p>Bodt, Timotheus Adrianus. 2024. <em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><span><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&code=view&bookID=146">LANGUAGE AND LINGUISTICS > (sinica.edu.tw)</a></span></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>
Sartang Lexicon - Khoina
<p>This dataset contains the sound files (original and cut) and transcribed sound files of Sartang - Khoina variety.</p> <p>Bodt, Timotheus Adrianus. 2024. <em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><span><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&code=view&bookID=146">LANGUAGE AND LINGUISTICS > (sinica.edu.tw)</a></span></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>
Khispi Lexicon
<p>These files contain the original sound files, the cut sound files, and the transcriptions of the lexicon of Khispi collected from Lish.</p> <p>Bodt, Timotheus Adrianus. 2024. <em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><span><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&code=view&bookID=146">LANGUAGE AND LINGUISTICS > (sinica.edu.tw)</a></span></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>
Pere lexicon of flora and fauna
<p>A lexicon of flora and faun terms for Pere (Bere, Mbre) of Côte d’Ivoire. Pere (in the literature also spelled Pɛrɛ, Bere, Mbre) is a seriously endangered language of central Côte d’Ivoire. It is listed as “Mbre” in Glottolog (viewed July 2018), code mbre1244, and in ISO 639-3, code mka.</p>
Comprehensive understanding of Tn5 insertion preference recovers expansive transcription regulatory lexicon
<p>This repository stores pre-calculated mappability files for BiasFreeATAC correction pipeline.</p>
Lexicon and example extensions from paper "Lexicon-based comments-oriented news sentiment analyzer system"
<p>Lexicon and example extensions for the article "Moreo, Alejandro, et al. "Lexicon-based comments-oriented news sentiment analyzer system." <em>Expert Systems with Applications</em> 39.10 (2012): 9166-9180." Founded by Ministerio de Educación y Ciencia and Junta de Andalucía with Projects: TIN2007-60199, TIC2009-5011 and TIN2007-67984</p>
Lexicon from the "Mics in the ears" experimental procedure Lexique issu du dispositif expérimental « Des micros dans les oreilles »
<p>« Mics in the Ears » binaural experiment in Cairo (Egypt): Vincent Battesti & Nicolas Puig, social anthropologists, asked inhabitants of Cairo megapolis in Egypt to record the surrounding urban sounds during one of their daily journeys (without the researcher), equipped with binaural microphones and GPS device. Participants have recorded in different Cairo neighbourhoods, and are themselves from different generations, social and economic backgrounds, and different genders.</p> <p>We obtained accounts of the soundwalks in which participants describe and comment on the sounds heard while relistening to their recorded walk. This material comprises many pages of transcriptions in Arabic from which we have extracted elements related to the verbalization of their sound experience. We identified 600 entries divided into five classes: “nouns,” “sound sources,” “qualificatives/descriptives,” “actions/ verbs,” and “localizations of the sound event.” Together they form the lexicon of what can be called the “natural language of sounds” in Cairo. See https://vbat.org/article831</p>
Emoji Sentiment Lexicons
<p>These are the emoji sentiment lexica derived from valence scores from cooccurrence with sentiment-carrying messages. One lexicon is based on a Twitter corpus and contains Unicode emojis, the other is based on a collection of Twitch chat logs and mainly contains valence values for Twitch emotes.</p>
Supplementary Materials for 'Colexification patterns in Europe: A study of persistence and diffusibility in the lexicon, based on the Database of Crosslinguistic Colexifications (CLICS3)' (Linguistic Typology)
<p>The folder contains the data and scripts used for the article on 'Colexification patterns in Europe: A study of persistence and diffusibility in the lexicon, based on the Database of Crosslinguistic Colexifications (CLICS3)'</p>
Lexicon
<p>English-Greek Lexicon of Electronics </p>
Lexicon_V2
<p>English-Greek Lexicon of Electronics</p> <p>Contains more than 31.000 entries.</p>
Komnzo lexicon
<p>This dataset contains a pdf-export of the most up to date version of the Komnzo lexicon.</p> <p>The material was recorded by Christian Döhler as part of a language documentation project for his PhD. The project was located at the <a href="http://chl.anu.edu.au/">School of Culture, History and Language</a> at the <a href="http://anu.edu.au/">Australian National University, Canberra</a>. For the most part it was funded by the <a href="http://dobes.mpi.nl/">DOBES project</a> of the <a href="https://www.volkswagenstiftung.de/en/foundation">Volkswagen Foundation</a>.</p>
JULIELab/MEmoLon (Individual Lexicons)
<p>A more conveniently formatted version of the emotion lexicons presented in our <a href="https://aclanthology.org/2020.acl-main.112/">ACL 2020 paper</a>. Different from the <a href="https://doi.org/10.5281/zenodo.3756607">original record</a>, in this record here, each lexicon is compressed into a <em>seperate</em> ZIP-file so that it can be downloaded individually. This record only contains the 'MTL_grouped' version of the lexicons. (This version has been used in the majority of the experiments in the paper.)</p> <p>The lexicons cover 91 languages in total. Each file name contains the <a href="https://en.wikipedia.org/wiki/List_of_ISO_639-1_codes">ISO 639-1 code</a> of the respective language. For example, the English lexicon can be found in the file `en.tsv.zip`.</p>
Comparative lexicon in Milang, Kera'a, Tawrã and Proto-Tani, with selected PTB or meso-level comparative reconstructions
<p>This data table forms the Appendix to Modi, Yankee (under review 2022). 'A branch through the fog: Milang, Tawrã and Kera’a.' This description will be updated after hopeful publication.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.