Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
85
datasets available to search
ShareScore release 0.9.0
Dataset results
85 results for “lexicon”
Senti-Pol-sr Sentiment Lexicon
Open the record for dataset details and reuse information.
Data from: Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types
<p>Concerns about gender bias in word embedding models have captured substantial attention in the algorithmic bias research literature. Other bias types however have received lesser amounts of scrutiny. This work describes a large-scale analysis of sentiment associations in popular word embedding models along the lines of gender and ethnicity but also along the less frequently studied dimensions of socioeconomic status, age, physical appearance, sexual orientation, religious sentiment and political leanings. Consistent with previous scholarly literature, this work has found systemic bias against given names popular among African-Americans in most embedding models examined. Gender bias in embedding models however appears to be multifaceted and often reversed in polarity to what has been regularly reported. Interestingly, using the common operationalization of the term bias in the fairness literature, novel types of so far unreported bias types in word embedding models have also been identified. Specifically, the popular embedding models analyzed here display negative biases against middle and working-class socioeconomic status, male children, senior citizens, plain physical appearance and intellectual phenomena such as Islamic religious faith, non-religiosity and conservative political orientation. Reasons for the paradoxical underreporting of these bias types in the relevant literature are probably manifold but widely held blind spots when searching for algorithmic bias and a lack of widespread technical jargon to unambiguously describe a variety of algorithmic associations could conceivably be playing a role. The causal origins for the multiplicity of loaded associations attached to distinct demographic groups within embedding models are often unclear but the heterogeneity of said associations and their potential multifactorial roots raises doubts about the validity of grouping them all under the umbrella term bias. Richer and more fine-grained terminology as well as a more comprehensive exploration of the bias landscape could help the fairness epistemic community to characterize and neutralize algorithmic discrimination more efficiently.</p>
Assessing the size of non-Māori-speakers' active Māori lexicon
<p>Initial release of data and code</p>
Lubrang Brokpa Lexicon - overview file
<p>These sound files constitute the elicitation of the lexical entries of the Basic Word List in the Lubrang variety of Brokpa. Lubrang village is a recent (early 20th century) settlement of Brokpa speaker originating in Sakteng village of Bhutan. They settled in the then Tibetan-administered area on land belonging to the Khispi people of Lish village, to which they continue to pay an annual tax. Lubrang Brokpa should hence be close to Merak and Sakteng (Bhutan) Brokpa, and not so close to Nyukmadung and Senge Brokpa spoken closer by. Because of the speaker’s paternal background there may be some admixture with Dirang Tshangla.</p> <p>Bodt, Timotheus Adrianus. 2024. <em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&code=view&bookID=146">LANGUAGE AND LINGUISTICS > (sinica.edu.tw)</a></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>
Data from: Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types
Open the record for dataset details and reuse information.
Supplementary materials for Gast, V. and M. Koptjevskaja-Tamm, 'Colexification patterns in Europe: A study of persistence and diffusibility in the lexicon, based on the Database of Crosslinguistic Colexifications (CLICS3)
<p>The folder contains the data, scripts and plots for the paper. It includes the following third party material (Open Access) in the folder DataStageI:</p> <p>1) the clics.sqlite database from https://github.com/clics/clics3; cf. Rzymski, Christoph and Tresoldi, Tiago et al. 2019. The Database of Cross-Linguistic Colexifications, reproducible analysis of cross- linguistic polysemies. <a href="https://doi.org/10.1038/s41597-019-0341-x">DOI: 10.1038/s41597-019-0341-x</a><br> 2) the files ccCosineDist.csv and pmiWorld.csv from Jäger (2018), 'Global-scale phylogenetic linguistic inference from lexical resources', Scientific Data 5, Article number: 180189<br> 3) languoid.csv from https://glottolog.org/meta/downloads; cf. Hammarström, Harald & Forkel, Robert & Haspelmath, Martin & Bank, Sebastian. 2020. Glottolog 4.2.1. Jena: Max Planck Institute for the Science of Human History.<br> (Available online at http://glottolog.org, Accessed on 2020-07-11.)<br> * asjp.tsv from https://asjp.clld.org/, cf. Wichmann, Søren, Eric W. Holman, and Cecil H. Brown (eds.). 2018. The ASJP Database (version 18). (current version: 2020)</p>
Data from: Lexicon-enhanced sentiment analysis framework using rule-based classification scheme
With the rapid increase in social networks and blogs, the social media services are increasingly being used by online communities to share their views and experiences about a particular product, policy and event. Due to economic importance of these reviews, there is growing trend of writing user reviews to promote a product. Nowadays, users prefer online blogs and review sites to purchase products. Therefore, user reviews are considered as an important source of information in Sentiment Analysis (SA) applications for decision making. In this work, we exploit the wealth of user reviews, available through the online forums, to analyze the semantic orientation of words by categorizing them into +ive and -ive classes to identify and classify emoticons, modifiers, general-purpose and domain-specific words expressed in the public's feedback about the products. However, the un-supervised learning approach employed in previous studies is becoming less efficient due to data sparseness, low accuracy due to non-consideration of emoticons, modifiers, and presence of domain specific words, as they may result in inaccurate classification of users' reviews. Lexicon-enhanced sentiment analysis based on Rule-based classification scheme is an alternative approach for improving sentiment classification of users' reviews in online communities. In addition to the sentiment terms used in general purpose sentiment analysis, we integrate emoticons, modifiers and domain specific terms to analyze the reviews posted in online communities. To test the effectiveness of the proposed method, we considered users reviews in three domains. The results obtained from different experiments demonstrate that the proposed method overcomes limitations of previous methods and the performance of the sentiment analysis is improved after considering emoticons, modifiers, negations, and domain specific terms when compared to baseline methods.
LEXICON OF THE UZBEK LANGUAGE AND THE PROCESS OF WORD ACQUISITION IN IT
Open the record for dataset details and reuse information.
SEMANTIC STUDY OF ETHNOGRAPHIC LEXICON
Open the record for dataset details and reuse information.
Estonian Pandemic Lexicon
Open the record for dataset details and reuse information.
Data from: Lexicon-enhanced sentiment analysis framework using rule-based classification scheme
Open the record for dataset details and reuse information.
A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation [ATAC-Seq]
GEO Series GSE188398. Homo sapiens. 18 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation [HiChIP-Seq]
GEO Series GSE188401. Homo sapiens. 17 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation
GEO Series GSE188405. synthetic construct; Homo sapiens. 90 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Other.
A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation [MPRA]
GEO Series GSE188403. Homo sapiens; synthetic construct. 19 samples. Type: Expression profiling by high throughput sequencing; Other.
Appetite Lexicon Training
ClinicalTrials.gov study NCT04576585. IPD Sharing: Not stated. Countries: 1. Publications: 0.
The dynamic, combinatorial cis-regulatory lexicon of epidermal differentiation
GEO Series GSE181416. Homo sapiens; Escherichia coli. 46 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Other.
A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation [RNA-Seq]
GEO Series GSE186947. Homo sapiens. 36 samples. Type: Expression profiling by high throughput sequencing.
The dynamic, combinatorial cis-regulatory lexicon of epidermal differentiation [PAS-seq]
GEO Series GSE181415. Homo sapiens. 26 samples. Type: Expression profiling by high throughput sequencing.
The dynamic, combinatorial cis-regulatory lexicon of epidermal differentiation [ChIP-seq]
GEO Series GSE181407. Homo sapiens. 0 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Third-party reanalysis.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.