Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

74

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

74 results for “sentiment analysis”

Learn how ShareScore rates datasets ↗
zenodo48/100

Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text

<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>

opencc-by-4.0May 2020View details →
zenodo48/100

A Sentiment Analysis Dataset for Code-Mixed Malayalam-English

<p>There is an increasing demand for sentiment analysis of text from social media which are mostly code-mixed. Systems trained on monolingual data fail for code-mixed data due to the complexity of mixing at different levels of the text. However, very few resources are available for code-mixed data to create models specific for this data. Although much research in multilingual and cross-lingual sentiment analysis has used semi-supervised or unsupervised methods, supervised methods still performs better. Only a few datasets for popular languages such as English-Spanish, English-Hindi, and English-Chinese are available. There are no resources available for Malayalam-English code-mixed data. This paper presents a new gold standard corpus for sentiment analysis of code-mixed text in Malayalam-English annotated by voluntary annotators. This gold standard corpus obtained a Krippendorff&rsquo;s alpha above 0.8 for the dataset. We use this new corpus to provide the benchmark for sentiment analysis in Malayalam-English code-mixed texts.</p>

opencc-by-4.0May 2020View details →
zenodo48/100

Large Dataset of Nigeria Covid-19 Tweets for Sentiment Analysis and Opinion Mining Tasks

<p><strong>Background</strong></p> <p>Information is essential for growth; without it, little can be accomplished. Data gathering has seen significant changes throughout the previous few centuries because of certain transitory medium. The look and style of information transference are affected by the employment of new and emerging technologies, some of which are efficient, others are reliable, and many more are quick and effective, but a few were disappointing for various reasons.</p> <p><strong>Aims</strong></p> <p>This study aims at using TextBlob and VADER analyser with historical tweets, to analyse emotional responses to the corona virus pandemic (covid-19). It shows us how much of a sociological, environmental, and economic impact it has in Nigeria, among other things. This study would be a tremendous step forward for students, researchers, and scholars who want to advance in fields like data science, machine learning, and deep learning.</p> <p><strong>Methodology</strong></p> <p>The hashtag &lsquo;covid-19&#39; was used to collect 1,048,575 tweets from Twitter. The tweets were pre-processed with a twitter tokenizer, and Valence Aware Dictionary for Sentiment Reasoning (VADER) and TextBlob were used for sentiment and text mining, respectively. Topic modelling was done with Latent Dirichlet Allocation (LDA). The simulated subjects, on the other hand, were visualized using Multidimensional scaling (MDS).</p> <p><strong>Results</strong></p> <p>The result of the VADER sentiment returned 39.8%, 31.3% and 28.9%, positive, neutral, and negative sentiment respectively while the result of the TextBlob sentiment returned 46.0%, 36.7% and 17.3%, neutral, positive, and negative sentiment, respectively.</p> <p><strong>Conclusion</strong></p> <p>With all of this, information from social media may be used to help organizations, governments, and nations around the world make smart and effective decisions about how to restrict and limit the negative effects of covid-19. Also know the opinion and challenges of people, then deal with problem of misinformation.</p> <p>It is concluded that with popular belief a significant number of the populace regards covid-19 as a virus that has come to stay, some believe it will eventually be conquered.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

MAVIS Twitter dataset: A collection of tweets and sentiment analysis in Spanish about vaccines and diseases during the period 2015-2018

<p>MAVIS dataset comprises a full knowledge base regarding Twitter messages published in Spanish during the period 2015-2018, in the context of sentiment analysis of specific vaccines and their related diseases. Such diseases and vaccines are summarized as follows:</p> <ul> <li>Invasive meningococcal disease (&ldquo;EMI&rdquo; in Spanish): Bexsero, Trumenba, Nimenrix</li> <li>Invasive pneumococcal disease (&ldquo;ENI&rdquo; in Spanish)</li> <li>Influenza</li> <li>Hepatitis</li> <li>Rotavirus: Rotarix, Rotateq</li> <li>Measles (&ldquo;Sarampi&oacute;n&rdquo; in Spanish) and MMR (&ldquo;Triple v&iacute;rica&rdquo; in Spanish)</li> <li>Sepsis</li> <li>Whooping cough (&ldquo;Tosferina&rdquo; in Spanish)</li> <li>Chickenpox (&ldquo;Varicela&rdquo; in Spanish): Varivax, Varilrix; and Shingles (&ldquo;Zoster&rdquo; in Spanish)</li> <li>Human papillomavirus infection (&ldquo;VPH&rdquo; in Spanish): Cervarix, Gardasil</li> </ul> <p>Tweets have been manually classified as having a negative or non-negative sentiment by 5 experts. Moreover, an automatic classification has been performed by 3 different tools: IBM Watson (now Watson Tone Analyzer, <a href="https://www.ibm.com/watson/services/tone-analyzer/">https://www.ibm.com/watson/services/tone-analyzer/</a>), Google Cloud Natural Language (<a href="https://cloud.google.com/natural-language">https://cloud.google.com/natural-language</a>), and Meaning Cloud (<a href="https://www.meaningcloud.com/">https://www.meaningcloud.com/</a>). IBM Watson and Google Cloud Natural Language returned a numerical sentiment score ranging from -1 to 1, while Meaning Cloud returned a categorical variable with the values &lsquo;P+&rsquo;, &lsquo;P&rsquo;, &lsquo;NEU&rsquo;, &lsquo;N&rsquo; and &lsquo;N+&rsquo;, which were converted to 1, 2, 3, 4 and 5 respectively.</p> <p>With these variables (IBM Watson, Google Cloud Natural Language, and Meaning Cloud annotations and the experts&rsquo; classification as the target label), a machine learning metamodel was developed. Tweets were also annotated with the sentiment output given by this classifier. &nbsp;&nbsp;</p> <p>The provided data includes intrinsic tweets information, intrinsic information regarding the users that posted the tweets, the keywords mentioned in each tweet, and the annotations that the experts, the tools, and the model gave to each tweet.</p> <p><strong>Funding</strong>: This dataset was obtained with funding from&nbsp;MSD, Spain under MAVIS Study (VEAP ID: 7789).</p> <p><strong>Current studies using this dataset at the moment of the publication</strong>:</p> <ul> <li>Rodr&iacute;guez-Gonz&aacute;lez et al., &ldquo;Creating a metamodel based on machine learning to identify the sentiment of vaccine and disease-related messages in Twitter: the MAVIS study&rdquo; in 2020 IEEE 33st International Symposium on Computer-Based Medical Systems (CBMS), Jul. 2020, p. 6. DOI: 10.1109/CBMS49503.2020.00053</li> <li>Rodr&iacute;guez-Gonz&aacute;lez et al., &quot;Identifying Polarity in Tweets from an Imbalanced Dataset about Diseases and Vaccines Using a Meta-Model Based on Machine Learning Techniques&quot; in Applied Sciences, 2020, 10. DOI: 10.3390/app10249019</li> </ul>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Sentiment analysis in Galaxy with IMDB movie review dataset

<p>IMDB movie review sentiment classification dataset (Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. (2011).&nbsp;Learning Word Vectors for Sentiment Analysis.&nbsp;The 49th Annual Meeting of the Association for Computational Linguistics (ACL 2011)). For more information&nbsp;please refer to:&nbsp;https://ai.stanford.edu/~amaas/data/sentiment/<br> <br> The IMDB dataset was modified as follows to prepare it for use in a Galaxy Training Tutorial (https://training.galaxyproject.org/):<br> <br> The top 50 words are excluded (mostly stop words). Included&nbsp;the next 10,000 top words. Reviews are limited to&nbsp;500 words max (Longer reviews trimmed and shorter reviews are padded). 25,000 reviews are used for training and testing each. Files are&nbsp;in tsv (tab separated value) format to be consumed by Galaxy (www.usegalaxy.org).&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Sentiment analysis of tech media articles using VADER package and co-occurrence analysis during the COVID-19 pandemic (01.2020-06.2020)

<p><strong>Sources:&nbsp;</strong></p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe&nbsp;</li> <li>IEEE Spectrum&nbsp;</li> <li>Techforge&nbsp;</li> <li>Fastcompany&nbsp;</li> <li>The Guardian (Tech)&nbsp;</li> <li>Arstechnica&nbsp;</li> <li>Reuters&nbsp;</li> <li>Gizmodo&nbsp;</li> <li>ZDNet&nbsp;</li> <li>The Register&nbsp;</li> <li>The Verge&nbsp;</li> <li>TechCrunch&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood&#39;s scores would be positive, but the negative term would bring the paragraph&#39;s score down.</p> <p>The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. &amp; Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Dataset: Systematic Mapping Study on the Development and Application of Sentiment Analysis Tools in Software Engineering

<p>Update: We updated the data set in March 2022 by adding newly published papers and by providing more insights on how we analyzed them. Details can be found in the file &quot; SEnti-SMS.xlsx&quot;.</p> <p>----------</p> <p>Update: The updated version (-v2) contains the results of one more snowballing iteration and extracted information on the accuracy of the used methods.</p> <p>----------</p> <p>In 2020, we conducted a systematic literature review to explore the development and application of sentiment analysis tools in software engineering.</p> <p>Information on the execution of the SLR, its scope, the search string, etc. are presented in the paper linked below.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Dataset for sentiment analysis in Spanish

<p>This dataset is automatically generated by webscraping from sites such as Tripadvisor or Google Maps reviews. In these sites, the users post comments with ratings, allowing us to have tagged data. The code that generated this dataset can be found at the following URL:</p> <p><a href="https://github.com/fjramirezv/sentiment-webscraping">https://github.com/fjramirezv/sentiment-webscraping</a></p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

<p>We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria&mdash;Hausa, Igbo, Nigerian-Pidgin, and Yor&ugrave;b&aacute;&mdash;consisting of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), including a significant fraction of code-mixed tweets.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Dataset: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter

<p>This is the dataset, trained model, and software companion for the paper titled: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter accepted for the Workshop on Data for the Wellbeing of Most Vulnerable of the&nbsp;ICWSM 2022 conference.</p> <p>The COVID-19 pandemic has shown a measurable increase in the usage of sinophobic comments or terms on online social media platforms. In the United States, Asian Americans have been primarily targeted by violence and hate speech stemming from negative sentiments about the origins of the novel SARS-CoV-2 virus. While most published research focuses on extracting these sentiments from social media data, it does not connect the specific news events during the pandemic with changes in negative sentiment on social media platforms. In this work we combine and enhance publicly available resources with our own manually annotated set of tweets to create machine learning classification models to characterize the sinophobic behavior. We then applied our classifier to a pre-filtered longitudinal dataset spanning two years of pandemic related tweets and overlay our findings with relevant news events.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

CCUS Sentiment Analysis - Tweets Dataset

<p>The present dataset contains Tweets in any language supported by Twitter obtained during the months January to March 2023, with any mention to the topic CCS/CCUS. The scraping process were done in Python, using the official Twitter API. All tweets were manually annotated after being machine translated into English.</p> <p><strong>- Structure&nbsp;</strong><br>Every row contains:&nbsp;<br>1st cell (A): Language&nbsp;<br>2nd cell (B): Tweet-text&nbsp;<br>3rd cell (Cc: Benefit&nbsp;<br>4th cell (D): Concern&nbsp;<br>5th cell (E): Perception &ndash; Fight climate change&nbsp;<br>6th cell (F): Perception &ndash; Climate-friendly technology&nbsp;<br>7th cell (G): Perception &ndash; Extensive R&amp;D needed&nbsp;<br>8th cell (H): Perception &ndash; Better options than CCS&nbsp;<br>9th cell (I): Sentiment&nbsp;<br>10th cell (J): Relatedness&nbsp;<br>11th cell (K): Comments&nbsp;</p> <p><strong>- Annotations&nbsp;</strong><br><strong>Benefit&nbsp;</strong><br>Preventing c. change&nbsp;<br>Reducing c. change risks&nbsp;<br>Safeguarding jobs&nbsp;<br>Creating new jobs&nbsp;<br>Fossil energy production envir. friendly&nbsp;<br>Products envir. friendly&nbsp;<br>Reducing envir. impact&nbsp;<br>Other&nbsp;<br>None&nbsp;<br><strong>Concern&nbsp;</strong><br>Accidents&nbsp;<br>Leakages&nbsp;<br>Environmental&nbsp;<br>Earthquake-related&nbsp;<br>Increased local traffic&nbsp;<br>Investment&nbsp;<br>Greenwashing&nbsp;<br>Lock-in effects for fossil energy&nbsp;<br>Increase cost&nbsp;<br>Other&nbsp;<br>None&nbsp;<br><strong>Perception (Yes / No / None)&nbsp;</strong><br>Fight climate change&nbsp;<br>Climate-friendly technology&nbsp;<br>Extensive R&amp;D needed&nbsp;<br>Better options than CCS&nbsp;<br><strong>Sentiment&nbsp;</strong><br>Positive&nbsp;<br>Negative&nbsp;<br>Neutral&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Experimental datasets for sentiment analysis and emotion mining - Emotion Mining Toolkit (EMTk)

<p><strong>Description</strong></p> <p>Datasets for sentiment analysis and emotion mining, distributed with the Emotion Mining Toolkit (EMTk) Docker container (see <a href="https://collab-uniba.github.io/EMTk">https://collab-uniba.github.io/EMTk</a> for more):</p> <ul> <li>Stack Overflow - A couple of gold standards of 4,000+ posts, manually annotated for mining both emotions and polarity.</li> <li>Jira - A gold standard of ~4,000 issues, manually annotated for emotions.</li> </ul> <p><strong>Citation</strong></p> <p>Please, see the references below for the papers to cite. Do not cite this Zenodo upload directly.</p>

openother-openFeb 2019View details →
zenodo44/100

Sentiment analysis of tech media articles using VADER package and co-occurrence analysis

<p><strong>Sentiment analysis of tech media articles using VADER package and co-occurrence analysis</strong></p> <p><strong>Sources</strong>: Above 140k articles (01.2016-03.2019):</p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood&#39;s scores would be positive, but the negative term would bring the paragraph&#39;s score down.</p> <p>The presented tables include the most extreme co-occurring terms for the analysed social issue. The examples are chosen from the list of words with 30 most positive and 30 most negative sentiment. The presented graphs show the evolution of sentiments for social issues. The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. &amp; Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p> <p>&nbsp;</p> <p><strong>Files</strong></p> <p>sentiments_mod11.csv sentiment score based on chosen unigrams</p> <p>sentiments_mod22.csv sentiment score based on chosen bigrams</p> <p>sentiments_cooc_mod11.csv, sentiments_cooc_mod12.csv, sentiments_cooc_mod21.csv, sentiments_cooc_mod22.csv combinations of co-occurrences: unigrams-unigrams, unigrams-bigrams, bigrams-unigrams, bigrams-bigrams</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

AWARE: Dataset for Aspect-Based Sentiment Analysis of Apps Reviews

<p>&nbsp;</p> <p><em><strong>The&nbsp;peer-reviewed paper of&nbsp;AWARE dataset is published in ASEW 2021, and can be accessed&nbsp;through: <a href="http://doi.org/10.1109/ASEW52652.2021.00049">http://doi.org/10.1109/ASEW52652.2021.00049</a>.&nbsp;Kindly cite this paper when using AWARE dataset.</strong></em></p> <p>&nbsp;</p> <p>Aspect-Based Sentiment Analysis (ABSA) aims to identify the opinion (sentiment) with respect to a specific aspect. Since there is a lack of&nbsp;<em>smartphone apps reviews</em>&nbsp;dataset that is annotated to support the ABSA task, we present AWARE:&nbsp;<strong>A</strong>BSA&nbsp;<strong>W</strong>arehouse of&nbsp;<strong>A</strong>pps&nbsp;<strong>RE</strong>views.</p> <p>AWARE contains apps reviews from three different domains (Productivity, Social Networking, and Games), as each domain has its distinct functionalities and audience. Each sentence is annotated with three labels, as follows:&nbsp;</p> <ul> <li><strong>Aspect Term:&nbsp;</strong>a term that exists in the sentence and describes an aspect of the app that is expressed by the sentiment. A term value of &ldquo;N/A&rdquo; means that the term is not explicitly mentioned in the sentence.</li> <li><strong>Aspect Category:</strong>&nbsp;one of the pre-defined set of domain-specific categories that represent an aspect of the app (e.g., security, usability, etc.).</li> <li><strong>Sentiment:</strong>&nbsp;positive or negative.</li> </ul> <p><em>Note: games domain does not contain aspect terms.</em></p> <p>We provide a comprehensive dataset of 11323 sentences from the three domains, where each sentence is additionally annotated with a Boolean value indicating whether the sentence expresses a positive/negative opinion. In addition, we provide three separate datasets, one for each domain, containing only sentences that express opinions. The file named &ldquo;AWARE_metadata.csv&rdquo; contains a description of the dataset&rsquo;s columns.</p> <p><strong>How AWARE can be used?</strong></p> <p>We designed AWARE such that it can be used to serve various tasks. The tasks can be, but are not limited to:</p> <ul> <li>Sentiment Analysis.</li> <li>Aspect Term Extraction.</li> <li>Aspect Category Classification.</li> <li>Aspect Sentiment Analysis.</li> <li>Explicit/Implicit Aspect Term Classification.</li> <li>Opinion/Not-Opinion Classification.</li> </ul> <p>Furthermore, researchers can experiment with and investigate the effects of different domains on users&#39; feedback.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

BRI Sentiment Analysis

<p>This dataset contains news articles related to the Belt and Road Initiative (BRI) from various sources, collected between 2015&nbsp;and 2023 in English and from 2019 to 2023 in Chinese. The articles are labeled with sentiment scores for sentiment analysis, using a three-point scale: negative, neutral, positive&nbsp;The dataset aims to provide insights into the public perception of the BRI and its impact on various countries and regions. The data can be used for sentiment analysis, natural language processing, and machine learning research related to the BRI.</p> <p>Data source: media outlets in Chinese and in English (the name of the outlet is indicated in the column source)</p> <p>Sentiment analysis algorithms:</p> <p>- SnowNLP for Chinese (https://github.com/isnowfy/snownlp) - Scale 0-1<br> - Own algorithm based on FinABSA approach for English (https://github.com/guijinSON/FinABSA) - Scale -11</p> <p>The files contain the following columns</p> <ul> <li>title - title of the article</li> <li>date - date of the article</li> <li>source - name of the source</li> <li>month,</li> <li>year,</li> <li>quarter,</li> <li>sentiment - sentiment score calculated by the algorithms</li> <li>sentiment_label - sentiment label based on optimized thresholds (positive, neutral, negative)</li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Sentiment analysis of tech media articles using VADER package and co-occurrence analysis (01.2016-12.2019)

<p>Sentiment analysis of tech media articles using VADER package and co-occurrence analysis</p> <p>Sources with weights:</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5&nbsp;%</li> <li>IEEE Spectrum 5&nbsp;%</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood&#39;s scores would be positive, but the negative term would bring the paragraph&#39;s score down.</p> <p>&nbsp;</p> <p>*Hutto, C.J. &amp; Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Tunizi: Tunisian Arabizi Sentiment Analysis Dataset

<p>Tunizi is the first 100% Tunisian Arabizi sentiment analysis dataset. Tunisian Arabizi is the representation of the tunisian dialect written in Latin characters and numbers rather than Arabic letters.We gathered comments from social media platforms that express sentiment about popular topics. For this purpose, we extracted 100k comments using public streaming APIs.&nbsp; Tunizi was preprocessed by removing links, emoji symbols, and punctuations.</p> <p>The collected comments were manually annotated using an overall polarity: &nbsp; positive (1), negative (-1) and neutral (0) class. &nbsp; We divided the dataset into separate training, validation and test sets, with a ratio of 7:1:2&nbsp; with a balanced split where the number of comments from positive class and negative class are almost the same.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Datasets for paper: J. A. Cerón-Guzmán and E. León-Guzmán (2016), A Sentiment Analysis System of Spanish Tweets and Its Application in Colombia 2014 Presidential Election. SocialCom 2016. DOI: 10.1109/BDCloud-SocialCom-SustainCom.2016.47

<p>Datasets for paper: J. A. Cerón-Guzmán and E. León-Guzmán (2016), A Sentiment Analysis System of Spanish Tweets and Its Application in Colombia 2014 Presidential Election. The 9th IEEE International Conference on Social Computing and Networking (SocialCom). DOI: 10.1109/BDCloud-SocialCom-SustainCom.2016.47</p>

opencc-by-4.0Jul 2016View details →
zenodo40/100

VnEmoLex: A Vietnamese emotion lexicon for sentiment intensity analysis

<p>VnEmoLex is the moderate-sized data set annotated for eight basic emotions: joy, sadness, anger, fear, trust, disgust, surprise and anticipation for Vietnamese. It is built on the NRC Word-Emotion Association Lexicon (EmoLex)<sup>1 </sup>and the Viet Wordnet<sup>2</sup> . VnEmoLex has total 12,795 words of which 4431 words are from the EmoLex dictionary, 8364 words are taken from the Viet Wordnet.</p> <p><sup>1 </sup>http://saifmohammad.com/WebPages/NRC-Emotion-Lexicon.htm</p> <p><sup>2 </sup>http://http://viet.wordnet.vn/wnms/</p>

opencc-by-4.0May 2017View details →
zenodo40/100

Synthetic Product Desirability Datasets for Sentiment Analysis Testing

<p><strong>Overview:</strong><br>This collection contains three synthetic datasets produced by gpt-4o-mini for sentiment analysis and PDT (Product Desirability Toolkit) testing. Each dataset contains 1000 hypothetical software product reviews with the aim to produce a diversity of sentiment and text. The datasets were created as part of the research described in:</p> <p>Hastings, J.D., Weitl-Harms, S., Doty, J., Myers, Z. L., and Thompson, W., &ldquo;Utilizing Large Language Models to Synthesize Product Desirability Datasets,&rdquo; in Proceedings of the 2024 IEEE International Conference<br>on Big Data (BigData-24), Workshop on Large Language and Foundation Models (WLLFM-24), Dec. 2024.<br><a href="https://arxiv.org/abs/2411.13485">https://arxiv.org/abs/2411.13485</a>.</p> <p>Briefly, each row in the datasets was produced as follows:<br>1) Word+Review: The LLM selected a word and synthesized a review that would align with a random target sentiment.<br>2) Review+Word: The LLM produced a review to align with the target sentiment score, and then selected a word appropriate for the review.<br>3) Supply-Word: A word was supplied to the LLM which was then scored, and a review was produced to align with that score.</p> <p>For sentiment analysis and PDT testing, the two columns of main interest across the datasets are likely 'Selected Word' and 'Hypothetical Review'.</p> <p><strong>License:</strong><br>This data is licensed under the CC Attribution 4.0 international license, and may be taken and used freely with credit given. Cite as:</p> <p>Hastings, J., Weitl-Harms, S., Doty, J., Myers, Z., &amp; Thompson, W. (2024). Synthetic Product Desirability Datasets for Sentiment Analysis Testing (1.0.0). Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.14188456" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.14188456</a></p>

opencc-by-4.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record