Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

5 results for “code-mixing”

Learn how ShareScore rates datasets ↗
zenodo48/100

Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text

<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>

opencc-by-4.0May 2020View details →
zenodo48/100

A Sentiment Analysis Dataset for Code-Mixed Malayalam-English

<p>There is an increasing demand for sentiment analysis of text from social media which are mostly code-mixed. Systems trained on monolingual data fail for code-mixed data due to the complexity of mixing at different levels of the text. However, very few resources are available for code-mixed data to create models specific for this data. Although much research in multilingual and cross-lingual sentiment analysis has used semi-supervised or unsupervised methods, supervised methods still performs better. Only a few datasets for popular languages such as English-Spanish, English-Hindi, and English-Chinese are available. There are no resources available for Malayalam-English code-mixed data. This paper presents a new gold standard corpus for sentiment analysis of code-mixed text in Malayalam-English annotated by voluntary annotators. This gold standard corpus obtained a Krippendorff&rsquo;s alpha above 0.8 for the dataset. We use this new corpus to provide the benchmark for sentiment analysis in Malayalam-English code-mixed texts.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Code-mixed Indonesian-Javanese-English Twitter Dataset

<p>This is a Twitter dataset for code-mixed language identification. The dataset contains mixed Indonesian, Javanese, and English words.&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Annotated Dataset for Bilingual Code-Mixed English-Malay Sentiment Analysis and Sarcasm Detection in Public Security Domain

<p>Tweets from X, and post with comment from TikTok was acquired <span>from 11 September until 21 September 2022</span>. Data from both platforms was merged and selected. Three annotators manually labelling the selected data for sentiment and sarcasm. Sentiment labels are &lsquo;positive&rsquo;, &lsquo;negative&rsquo;, and &lsquo;neutral&rsquo;. Sarcasm label is &lsquo;sarcastic&rsquo; and &lsquo;not sarcastic&rsquo;. Majority voting is considered for each label. Language identification label produced for each data.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

SemEval-2020 Task 9: Overview of Sentiment Analysis of Code-Mixed Tweets

<p>There are 2 sub-tasks: sentiment analysis for Spanglish (Spanish-English) and for Hinglish (Hindi-English).</p> <p>The sentiment classes are Positive, negative, neutral.&nbsp;</p> <p>Hinglish dataset has 20k instances.</p> <p>Spanglish dataset has ~19k instances.&nbsp;</p> <p>Website:&nbsp;<a href="https://ritual-uh.github.io/sentimix2020/">https://ritual-uh.github.io/sentimix2020/</a></p>

opencc-by-4.0Aug 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record