Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

85

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

85 results for “lexicon”

Learn how ShareScore rates datasets ↗
zenodo32/100

Senti-Pol-sr Sentiment Lexicon

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
dryad32/100

Data from: Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types

<p>Concerns about gender bias in word embedding models have captured substantial attention in the algorithmic bias research literature. Other bias types however have received lesser amounts of scrutiny. This work describes a large-scale analysis of sentiment associations in popular word embedding models along the lines of gender and ethnicity but also along the less frequently studied dimensions of socioeconomic status, age, physical appearance, sexual orientation, religious sentiment and political leanings. Consistent with previous scholarly literature, this work has found systemic bias against given names popular among African-Americans in most embedding models examined. Gender bias in embedding models however appears to be multifaceted and often reversed in polarity to what has been regularly reported. Interestingly, using the common operationalization of the term bias in the fairness literature, novel types of so far unreported bias types in word embedding models have also been identified. Specifically, the popular embedding models analyzed here display negative biases against middle and working-class socioeconomic status, male children, senior citizens, plain physical appearance and intellectual phenomena such as Islamic religious faith, non-religiosity and conservative political orientation. Reasons for the paradoxical underreporting of these bias types in the relevant literature are probably manifold but widely held blind spots when searching for algorithmic bias and a lack of widespread technical jargon to unambiguously describe a variety of algorithmic associations could conceivably be playing a role. The causal origins for the multiplicity of loaded associations attached to distinct demographic groups within embedding models are often unclear but the heterogeneity of said associations and their potential multifactorial roots raises doubts about the validity of grouping them all under the umbrella term bias. Richer and more fine-grained terminology as well as a more comprehensive exploration of the bias landscape could help the fairness epistemic community to characterize and neutralize algorithmic discrimination more efficiently.</p>

opencc-zeroApr 2020View details →
zenodo32/100

Assessing the size of non-Māori-speakers' active Māori lexicon

<p>Initial release of data and code</p>

openother-openJul 2023View details →
zenodo32/100

Lubrang Brokpa Lexicon - overview file

<p>These sound files constitute the elicitation of the lexical entries of the Basic Word List in the Lubrang variety of Brokpa. Lubrang village is a recent (early 20th century) settlement of Brokpa speaker originating in Sakteng village of Bhutan. They settled in the then Tibetan-administered area on land belonging to the Khispi people of Lish village, to which they continue to pay an annual tax. Lubrang Brokpa should hence be close to Merak and Sakteng (Bhutan) Brokpa, and not so close to Nyukmadung and Senge Brokpa spoken closer by. Because of the speaker&rsquo;s paternal background there may be some admixture with Dirang Tshangla.</p> <p>Bodt, Timotheus Adrianus. 2024.&nbsp;<em>Proto-Western Kho-Bwa: Reconstructing a communities' past through language.</em> Academia Sinica Languages and Linguistics monograph series number 67. Taipei: Academia Sinica.</p> <p><a href="https://www.ling.sinica.edu.tw/item/en?act=publish_book&amp;code=view&amp;bookID=146">LANGUAGE AND LINGUISTICS &gt; (sinica.edu.tw)</a></p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collector&nbsp;of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>

opencc-by-4.0Aug 2024View details →
dryad32/100

Data from: Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types

Open the record for dataset details and reuse information.

publicApr 2020View details →
zenodo28/100

Supplementary materials for Gast, V. and M. Koptjevskaja-Tamm, 'Colexification patterns in Europe: A study of persistence and diffusibility in the lexicon, based on the Database of Crosslinguistic Colexifications (CLICS3)

<p>The folder contains the data, scripts and plots for the paper. It includes the following third party material (Open Access) in the folder DataStageI:</p> <p>1) the clics.sqlite database from https://github.com/clics/clics3; cf. Rzymski, Christoph and Tresoldi, Tiago et al. 2019. The Database of Cross-Linguistic Colexifications, reproducible analysis of cross- linguistic polysemies. <a href="https://doi.org/10.1038/s41597-019-0341-x">DOI: 10.1038/s41597-019-0341-x</a><br> 2) the files ccCosineDist.csv and pmiWorld.csv from J&auml;ger (2018), &#39;Global-scale phylogenetic linguistic inference from lexical resources&#39;, Scientific Data 5, Article&nbsp;number:&nbsp;180189<br> 3) languoid.csv from https://glottolog.org/meta/downloads; cf. Hammarstr&ouml;m, Harald &amp; Forkel, Robert &amp; Haspelmath, Martin &amp; Bank, Sebastian. 2020. Glottolog 4.2.1. Jena: Max Planck Institute for the Science of Human History.<br> (Available online at http://glottolog.org, Accessed on 2020-07-11.)<br> * asjp.tsv from https://asjp.clld.org/, cf. Wichmann, S&oslash;ren, Eric W. Holman, and Cecil H. Brown (eds.). 2018. The ASJP Database (version 18). (current version: 2020)</p>

opencc-by-4.0Jul 2020View details →
dryad28/100

Data from: Lexicon-enhanced sentiment analysis framework using rule-based classification scheme

With the rapid increase in social networks and blogs, the social media services are increasingly being used by online communities to share their views and experiences about a particular product, policy and event. Due to economic importance of these reviews, there is growing trend of writing user reviews to promote a product. Nowadays, users prefer online blogs and review sites to purchase products. Therefore, user reviews are considered as an important source of information in Sentiment Analysis (SA) applications for decision making. In this work, we exploit the wealth of user reviews, available through the online forums, to analyze the semantic orientation of words by categorizing them into +ive and -ive classes to identify and classify emoticons, modifiers, general-purpose and domain-specific words expressed in the public's feedback about the products. However, the un-supervised learning approach employed in previous studies is becoming less efficient due to data sparseness, low accuracy due to non-consideration of emoticons, modifiers, and presence of domain specific words, as they may result in inaccurate classification of users' reviews. Lexicon-enhanced sentiment analysis based on Rule-based classification scheme is an alternative approach for improving sentiment classification of users' reviews in online communities. In addition to the sentiment terms used in general purpose sentiment analysis, we integrate emoticons, modifiers and domain specific terms to analyze the reviews posted in online communities. To test the effectiveness of the proposed method, we considered users reviews in three domains. The results obtained from different experiments demonstrate that the proposed method overcomes limitations of previous methods and the performance of the sentiment analysis is improved after considering emoticons, modifiers, negations, and domain specific terms when compared to baseline methods.

opencc-zeroDec 2016View details →
zenodo28/100

LEXICON OF THE UZBEK LANGUAGE AND THE PROCESS OF WORD ACQUISITION IN IT

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

SEMANTIC STUDY OF ETHNOGRAPHIC LEXICON

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo28/100

Estonian Pandemic Lexicon

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
dryad28/100

Data from: Lexicon-enhanced sentiment analysis framework using rule-based classification scheme

Open the record for dataset details and reuse information.

publicFeb 2017View details →
geo24/100

A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation [ATAC-Seq]

GEO Series GSE188398. Homo sapiens. 18 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenAug 2022View details →
geo24/100

A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation [HiChIP-Seq]

GEO Series GSE188401. Homo sapiens. 17 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenAug 2022View details →
geo24/100

A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation

GEO Series GSE188405. synthetic construct; Homo sapiens. 90 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Other.

openGEO-OpenAug 2022View details →
geo24/100

A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation [MPRA]

GEO Series GSE188403. Homo sapiens; synthetic construct. 19 samples. Type: Expression profiling by high throughput sequencing; Other.

openGEO-OpenAug 2022View details →
ClinicalTrials.gov24/100

Appetite Lexicon Training

ClinicalTrials.gov study NCT04576585. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
geo24/100

The dynamic, combinatorial cis-regulatory lexicon of epidermal differentiation

GEO Series GSE181416. Homo sapiens; Escherichia coli. 46 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing; Other.

openGEO-OpenAug 2021View details →
geo24/100

A cis-regulatory lexicon of DNA motif combinations mediating cell type-specific gene regulation [RNA-Seq]

GEO Series GSE186947. Homo sapiens. 36 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2022View details →
geo20/100

The dynamic, combinatorial cis-regulatory lexicon of epidermal differentiation [PAS-seq]

GEO Series GSE181415. Homo sapiens. 26 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2021View details →
geo20/100

The dynamic, combinatorial cis-regulatory lexicon of epidermal differentiation [ChIP-seq]

GEO Series GSE181407. Homo sapiens. 0 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Third-party reanalysis.

openGEO-OpenAug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record