Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
141
datasets available to search
ShareScore release 0.9.0
Dataset results
141 results for “Sentiment”
Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text
<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>
A Sentiment Analysis Dataset for Code-Mixed Malayalam-English
<p>There is an increasing demand for sentiment analysis of text from social media which are mostly code-mixed. Systems trained on monolingual data fail for code-mixed data due to the complexity of mixing at different levels of the text. However, very few resources are available for code-mixed data to create models specific for this data. Although much research in multilingual and cross-lingual sentiment analysis has used semi-supervised or unsupervised methods, supervised methods still performs better. Only a few datasets for popular languages such as English-Spanish, English-Hindi, and English-Chinese are available. There are no resources available for Malayalam-English code-mixed data. This paper presents a new gold standard corpus for sentiment analysis of code-mixed text in Malayalam-English annotated by voluntary annotators. This gold standard corpus obtained a Krippendorff’s alpha above 0.8 for the dataset. We use this new corpus to provide the benchmark for sentiment analysis in Malayalam-English code-mixed texts.</p>
Large Dataset of Nigeria Covid-19 Tweets for Sentiment Analysis and Opinion Mining Tasks
<p><strong>Background</strong></p> <p>Information is essential for growth; without it, little can be accomplished. Data gathering has seen significant changes throughout the previous few centuries because of certain transitory medium. The look and style of information transference are affected by the employment of new and emerging technologies, some of which are efficient, others are reliable, and many more are quick and effective, but a few were disappointing for various reasons.</p> <p><strong>Aims</strong></p> <p>This study aims at using TextBlob and VADER analyser with historical tweets, to analyse emotional responses to the corona virus pandemic (covid-19). It shows us how much of a sociological, environmental, and economic impact it has in Nigeria, among other things. This study would be a tremendous step forward for students, researchers, and scholars who want to advance in fields like data science, machine learning, and deep learning.</p> <p><strong>Methodology</strong></p> <p>The hashtag ‘covid-19' was used to collect 1,048,575 tweets from Twitter. The tweets were pre-processed with a twitter tokenizer, and Valence Aware Dictionary for Sentiment Reasoning (VADER) and TextBlob were used for sentiment and text mining, respectively. Topic modelling was done with Latent Dirichlet Allocation (LDA). The simulated subjects, on the other hand, were visualized using Multidimensional scaling (MDS).</p> <p><strong>Results</strong></p> <p>The result of the VADER sentiment returned 39.8%, 31.3% and 28.9%, positive, neutral, and negative sentiment respectively while the result of the TextBlob sentiment returned 46.0%, 36.7% and 17.3%, neutral, positive, and negative sentiment, respectively.</p> <p><strong>Conclusion</strong></p> <p>With all of this, information from social media may be used to help organizations, governments, and nations around the world make smart and effective decisions about how to restrict and limit the negative effects of covid-19. Also know the opinion and challenges of people, then deal with problem of misinformation.</p> <p>It is concluded that with popular belief a significant number of the populace regards covid-19 as a virus that has come to stay, some believe it will eventually be conquered.</p>
SenTopX: A Benchmark Twitter Dataset for User Sentiment on Various Topics
<p>This is a longitudinal Twitter dataset of 143K users during the period 2017-2021. The following is the detail of all the files:</p> <ul> <li><a href="11243662" target="_blank" rel="noopener noreferrer">SenTopX_userIDs.txt</a>: contains user IDs of 143K Twitter users.</li> <li><a href="../api/records/11243662/draft/files/userIDs_tweetIDs.zip/content" target="_blank" rel="noopener noreferrer">userIDs_tweetIDs.zip</a>: contains Tweet IDs of users, the name of the file is the user ID and the file contains the list of all the tweet IDs.</li> <li><a href="../api/records/11243662/draft/files/users_16_perspective_toxicity_scores.csv/content" target="_blank" rel="noopener noreferrer">users_16_perspective_toxicity_scores.csv</a> contains user IDs and 16 median Perspective API scores, the vector is shared as mean, median, and Gini Index of scores calculated over all tweets of a user.</li> <li><a href="../api/records/11243662/draft/files/LDAvis_top30_words_for_extracted_topics.csv/content" target="_blank" rel="noopener noreferrer">LDAvis_top30_words_for_extracted_topics.csv</a> contains the top 30 most relevant words extracted from each topic extracted by tweet-level topic modeling using the BERTweet topic model.</li> <li><a href="../api/records/11243662/draft/files/topic_modelling_statistics_per_user.csv/content" target="_blank" rel="noopener noreferrer">topic_modelling_statistics_per_user.csv</a> contains important and relevant statistics related to topic modeling results: <ul> <li> <p>1. user: This column represents the identifier for the user. Each row in the CSV corresponds to a specific user, and this column helps to track and differentiate between the users.</p> <p>2. avg_topic_probability: This column contains the average probability of the topics for each user calculated across all of the tweets in order to compare users in a meaningful way. It represents the average likelihood that a particular user discusses various topics over the observed period.</p> <p>3. maximum_topic_avg: This column holds the value of the highest average probability among all topics for each user. It indicates the topic that the user most frequently discusses, on average.</p> <p>4. index_max_avg_topic_probability_200: This column specifies the index or identifier of the topic with the highest average probability out of 200 possible topics. It shows which topic (out of 200) the user discusses the most.</p> <p>5. global_avg: This column includes the global average probability of topics across all users. It provides a baseline or overall average topic probability that can be used for comparative purposes.</p> <p>6. max_global_avg: This column contains the maximum global average probability across all topics for all users. It identifies the most discussed topic across the entire user base.</p> <p>7. index_max_global_avg: This column shows the index or identifier of the topic with the highest global average probability. It indicates which topic (out of 200) is the most popular across all users.</p> <p>8. entropy_200_topic: This column represents the entropy of the topics for each user, calculated over 200 topics. Entropy measures the diversity or unpredictability in the user's discussion of topics, with higher entropy indicating more varied topic discussion.</p> <p>In summary, these columns are used to analyze the topic engagement and preferences of users on a platform, highlighting the most frequently discussed topics, the variability in topic discussions, and how individual user behavior compares to overall trends.</p> </li> </ul> </li> </ul>
Sentiment polarity lexicon of Bosnian language
<p>First sentiment annotated lexicon of the Bosnian language.</p> <p>The lexicon is divided into two files: positive and negative polarity.</p> <p>The lists are prepared in separate files, each file populated by words of the aforementioned polarity listed each word in a separate row. </p> <p>The positive polarity list BOSNIAN_POSITIVE.txt holds 1219 words.</p> <p>The negative polarity list BOSNIAN_NEGATIVE.txt holds 3935 words.</p>
CED Sentiment Data
<p>1.Introduction</p> <p>This dashboard is aimed at providing a comprehensive data source that features the most important indicators related to the Chinese economy. The dashboard is divided into three sections, namely, sentiment data, high-frequency data, and low-frequency indicators. It will be updated on a monthly basis and the data can be downloaded directly from the webpage.</p> <p>General context</p> <p>The China Horizons project aims to contribute towards a better understanding of the Chinese economy and its impact on the global economy. The project seeks to achieve this goal by providing timely, accurate, and comprehensive data and analysis. The development of the dashboard is one of the tasks associated with this project, and it is expected to help researchers, policymakers, and other stakeholders to keep track of economic changes in China.</p> <p>2.Methodological approach</p> <p>China’s Economy Data (CED) Platform</p> <p>The China’s Economy Data (CED) Platform is a web application that has been developed using the Python programming language. This platform is available online and provides users with access to a wide range of data related to the Chinese economy.</p> <p>The CED platform has been designed to be user-friendly and intuitive, with a simple interface that makes it easy for users to navigate and access the data they need. It is a comprehensive platform that features data related to various aspects of the Chinese economy, such as consumer and investment sentiment, inflation, employment among others.</p> <p>The platform is built on a robust database that is regularly updated with the latest economic data from various sources, including government agencies and international organizations. This ensures that users have access to accurate and up-to-date information.</p> <p> </p> <p>Sentiment Data</p> <p>The aim of this section is to collect relevant data and analyze it in order to gauge the overall sentiment of consumers and investors towards the Chinese economy.</p> <p>The consumer sentiment data has been derived by tracking user conversations on social media platform Weibo and analyzing them using the SNOWNLP sentiment analysis algorithm. This data has been collected across three main consumer categories - food, automobile, and fashion industry - and has been cleaned to remove irrelevant messages, ads, and noise. The sentiment has been calculated using the SnowNLP algorithm.</p> <p>The investment sentiment data, on the other hand, has been calculated based on news articles published in the Chinese media and analyzed using the <a href="https://huggingface.co/uer/roberta-base-finetuned-jd-binary-chinese">https://huggingface.co/uer/roberta-base-finetuned-jd-binary-chinese</a> algorithm.</p> <p>In addition to these, the China Horizons dashboard also includes the China Economic Policy Uncertainty Index, which has been calculated following the methodology presented by Hong-Kong Baptist University (https://cbade.hkbu.edu.hk/epu-mainland-china/). This index provides an indication of the level of uncertainty in the Chinese economy, and is an important indicator for investors and policymakers</p> <p>High-frequency Data</p> <p> </p> <p>In this section, we collect and process fast-changing publicly available data from different data sources such as National Bureau of Statistics, The People's Bank of China China Foreign Exchange Trade System</p> <p>Low-frequency Data</p> <p> </p> <p>Similarly to the High Frequency data we collect the data that is publish quarterly or less often, and we transform it to the format required by the platform. Among low-frequency data sources, we can find:</p> <p>Chinese Ministry of Financel, State Administration of Foreign Exchange & National Bureau of Statistics,The People's Bank of China, China Ministry of Human Resources and Social Security, CEIC.</p> <p>3.Summary of activities and/or research findings</p> <p>China’s Economy Data (CED) Platform</p> <p> </p> <ul> <li> platform that includes a dashboard, a database</li> </ul> <p>Data Update Pipeline</p> <ul> <li>a pipeline that ingests and processes the input data and updates the dashboard</li> </ul> <p>4.Conclusions and future steps</p> <p>The platform gives a clear vision of key economic indicators in China. It will be updated once a month.</p> <p>5.Bibliographical references (if applicable)</p> <p><a href="https://cbade.hkbu.edu.hk/wp-content/uploads/2020/08/measuring_economic_policy_uncertainty_in_china_oct_2019.pdf">Huang, Y., and Luk, P. (2019). “Measuring Economic Policy Uncertainty in China.” <em>China Economic Review</em></a>.</p> <p>Zhou, Bingqian, Yuhan Zhu and Xiabing Mao. “Sentiment Analysis on Power Rationing Micro Blog Comments Based on SnowNLP-SVM-LDA Model.” Highlights in Science, Engineering and Technology (2022): n. pag.</p>
MAVIS Twitter dataset: A collection of tweets and sentiment analysis in Spanish about vaccines and diseases during the period 2015-2018
<p>MAVIS dataset comprises a full knowledge base regarding Twitter messages published in Spanish during the period 2015-2018, in the context of sentiment analysis of specific vaccines and their related diseases. Such diseases and vaccines are summarized as follows:</p> <ul> <li>Invasive meningococcal disease (“EMI” in Spanish): Bexsero, Trumenba, Nimenrix</li> <li>Invasive pneumococcal disease (“ENI” in Spanish)</li> <li>Influenza</li> <li>Hepatitis</li> <li>Rotavirus: Rotarix, Rotateq</li> <li>Measles (“Sarampión” in Spanish) and MMR (“Triple vírica” in Spanish)</li> <li>Sepsis</li> <li>Whooping cough (“Tosferina” in Spanish)</li> <li>Chickenpox (“Varicela” in Spanish): Varivax, Varilrix; and Shingles (“Zoster” in Spanish)</li> <li>Human papillomavirus infection (“VPH” in Spanish): Cervarix, Gardasil</li> </ul> <p>Tweets have been manually classified as having a negative or non-negative sentiment by 5 experts. Moreover, an automatic classification has been performed by 3 different tools: IBM Watson (now Watson Tone Analyzer, <a href="https://www.ibm.com/watson/services/tone-analyzer/">https://www.ibm.com/watson/services/tone-analyzer/</a>), Google Cloud Natural Language (<a href="https://cloud.google.com/natural-language">https://cloud.google.com/natural-language</a>), and Meaning Cloud (<a href="https://www.meaningcloud.com/">https://www.meaningcloud.com/</a>). IBM Watson and Google Cloud Natural Language returned a numerical sentiment score ranging from -1 to 1, while Meaning Cloud returned a categorical variable with the values ‘P+’, ‘P’, ‘NEU’, ‘N’ and ‘N+’, which were converted to 1, 2, 3, 4 and 5 respectively.</p> <p>With these variables (IBM Watson, Google Cloud Natural Language, and Meaning Cloud annotations and the experts’ classification as the target label), a machine learning metamodel was developed. Tweets were also annotated with the sentiment output given by this classifier. </p> <p>The provided data includes intrinsic tweets information, intrinsic information regarding the users that posted the tweets, the keywords mentioned in each tweet, and the annotations that the experts, the tools, and the model gave to each tweet.</p> <p><strong>Funding</strong>: This dataset was obtained with funding from MSD, Spain under MAVIS Study (VEAP ID: 7789).</p> <p><strong>Current studies using this dataset at the moment of the publication</strong>:</p> <ul> <li>Rodríguez-González et al., “Creating a metamodel based on machine learning to identify the sentiment of vaccine and disease-related messages in Twitter: the MAVIS study” in 2020 IEEE 33st International Symposium on Computer-Based Medical Systems (CBMS), Jul. 2020, p. 6. DOI: 10.1109/CBMS49503.2020.00053</li> <li>Rodríguez-González et al., "Identifying Polarity in Tweets from an Imbalanced Dataset about Diseases and Vaccines Using a Meta-Model Based on Machine Learning Techniques" in Applied Sciences, 2020, 10. DOI: 10.3390/app10249019</li> </ul>
Sentiment analysis in Galaxy with IMDB movie review dataset
<p>IMDB movie review sentiment classification dataset (Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. (2011). Learning Word Vectors for Sentiment Analysis. The 49th Annual Meeting of the Association for Computational Linguistics (ACL 2011)). For more information please refer to: https://ai.stanford.edu/~amaas/data/sentiment/<br> <br> The IMDB dataset was modified as follows to prepare it for use in a Galaxy Training Tutorial (https://training.galaxyproject.org/):<br> <br> The top 50 words are excluded (mostly stop words). Included the next 10,000 top words. Reviews are limited to 500 words max (Longer reviews trimmed and shorter reviews are padded). 25,000 reviews are used for training and testing each. Files are in tsv (tab separated value) format to be consumed by Galaxy (www.usegalaxy.org). </p>
AckSent: Human Annotated Dataset of Support and Sentiments in Dissertation Acknowledgments
<p>This data is supplementary to the paper:</p> <blockquote> <p><em>Manika Lamba, You Peng, Sophie Nikolov, and J. Stephen Downie. 2024. <strong>AckSent: Human Annotated Dataset of Support and Sentiments in Dissertation Acknowledgments</strong>. In The 2024 ACM/IEEE Joint Conference on Digital Libraries (JCDL ’24), December 2024, Hong Kong, China. ACM, New York, NY, USA, 5 pages. https://doi.org/10.1145/3677389.3702594</em></p> </blockquote>
Sentiment analysis of tech media articles using VADER package and co-occurrence analysis during the COVID-19 pandemic (01.2020-06.2020)
<p><strong>Sources: </strong></p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p> </p> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood's scores would be positive, but the negative term would bring the paragraph's score down.</p> <p>The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>
The ALPIN Sentiment Dictionary: Austrian Language Polarity in Newspapers
<p>These datasets are part of the submitted paper for the LREC2022 conference entitled: "The ALPIN Sentiment Dictionary: Austrian Language Polarity in Newspapers"</p> <p>The various data sources, as well as the methodology, are explained in detail in the research paper which will be available soon.</p> <p>ALPIN stands for Austrian Language Polarity in Newspapers. The dictionary consists of three different parts which were merged together:</p> <ul> <li>Austrian Media Corpus: AMC (AMC_v1.0.csv)</li> <li>STANDARD posts: STP (STP_v1.0.csv)</li> <li>Austriacisms: AUT (AUT_v1.0.csv)</li> </ul> <p>Austrian Media Corpus (AMC) (Ransmayr et al., 2017) & STANDARD posts (STP) (Schabus et al., 2017) rely on the SPLM algorithm as used in SentiDraw (Sharma & Dutta 2021). Austriacisms (AUT) was generated by using the Best-Worst scaling (BWS) (Kiritchenko and Mohammad, 2017b). The AUT list was collected from the “Variantenwörterbuch des Deutschen” (Ammon et al., 2016) (thereby only selecting those words that only surface in Austrian German and in no other variety of German) and an austriacism list of Wikipedia (https://de.wikipedia.org/wiki/Liste_von_Austriazismen).</p> <p>The scores are scaled to the interval [-1, 1] using the min-max-abs scaling, ranging from negative to positive.</p> <p>References:<br> Sharma, S. S., & Dutta, G. (2021). SentiDraw: Using star ratings of reviews to develop domain specific sentiment lexicon for polarity determination. Information Processing & Management, 58(1), 102412.<br> Kiritchenko, S. and Mohammad, S. M. (2017b). Capturing reliable fine-grained sentiment associations by crowdsourcing and best-worst scaling.<br> Schabus, D., Skowron, M., & Trapp, M. (2017). One Million Posts: A Data Set of German Online Discussions. Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, 1241–1244. https://doi.org/10.1145/3077136.3080711<br> Ransmayr, J., Mörth, K., & Ďurčo, M. (2017). AMC (Austrian Media Corpus). In Korpusbasierte Forschungen zum österreichischen Deutsch. In Digitale Methoden der Korpusforschung in Österreich (= Veröffentlichungen zur Linguistik und Kommunikationsforschung Nr. 30) (pp. 27–38). Verlag der Österreichischen Akademie der Wissenschaften.<br> Ammon, U., Bickel, H., & Ebner, J. (2016). Variantenwörterbuch des Deutschen : die Standardsprache in Österreich, der Schweiz, Deutschland, Liechtenstein, Luxemburg, Ostbelgien und Südtirol sowie Rumänien, Namibia und Mennonitensiedlungen. Walter de Gruyter.</p>
Dataset: Systematic Mapping Study on the Development and Application of Sentiment Analysis Tools in Software Engineering
<p>Update: We updated the data set in March 2022 by adding newly published papers and by providing more insights on how we analyzed them. Details can be found in the file " SEnti-SMS.xlsx".</p> <p>----------</p> <p>Update: The updated version (-v2) contains the results of one more snowballing iteration and extracted information on the accuracy of the used methods.</p> <p>----------</p> <p>In 2020, we conducted a systematic literature review to explore the development and application of sentiment analysis tools in software engineering.</p> <p>Information on the execution of the SLR, its scope, the search string, etc. are presented in the paper linked below.</p> <p> </p> <p> </p>
Dataset for sentiment analysis in Spanish
<p>This dataset is automatically generated by webscraping from sites such as Tripadvisor or Google Maps reviews. In these sites, the users post comments with ratings, allowing us to have tagged data. The code that generated this dataset can be found at the following URL:</p> <p><a href="https://github.com/fjramirezv/sentiment-webscraping">https://github.com/fjramirezv/sentiment-webscraping</a></p>
NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis
<p>We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria—Hausa, Igbo, Nigerian-Pidgin, and Yorùbá—consisting of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), including a significant fraction of code-mixed tweets.</p>
Dataset: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter
<p>This is the dataset, trained model, and software companion for the paper titled: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter accepted for the Workshop on Data for the Wellbeing of Most Vulnerable of the ICWSM 2022 conference.</p> <p>The COVID-19 pandemic has shown a measurable increase in the usage of sinophobic comments or terms on online social media platforms. In the United States, Asian Americans have been primarily targeted by violence and hate speech stemming from negative sentiments about the origins of the novel SARS-CoV-2 virus. While most published research focuses on extracting these sentiments from social media data, it does not connect the specific news events during the pandemic with changes in negative sentiment on social media platforms. In this work we combine and enhance publicly available resources with our own manually annotated set of tweets to create machine learning classification models to characterize the sinophobic behavior. We then applied our classifier to a pre-filtered longitudinal dataset spanning two years of pandemic related tweets and overlay our findings with relevant news events.</p>
Wikipedia Talk Page 'Climate Change' Sentiment and Toxicity Dataset
<p>The given dataset was prepared as part of a master's research project under the Master's program in Computational Social Systems at RWTH Aachen University.</p> <p>The talk page was parsed using the <a href="https://aclanthology.org/E17-3006/">GraWiTas tool</a> in JSON format. The file <em>Climate_change.comment_list.json </em> is raw export of discussions that needs to be cleaned before using it for calculating sentiment and toxicity scores.</p> <p>The sentiment scores were calculated using VADER and toxicity scores using the Perspective API by Google.</p> <p>Date & time of dump is: 27-06-2022 12:12 UTC+02:00</p>
CCUS Sentiment Analysis - Tweets Dataset
<p>The present dataset contains Tweets in any language supported by Twitter obtained during the months January to March 2023, with any mention to the topic CCS/CCUS. The scraping process were done in Python, using the official Twitter API. All tweets were manually annotated after being machine translated into English.</p> <p><strong>- Structure </strong><br>Every row contains: <br>1st cell (A): Language <br>2nd cell (B): Tweet-text <br>3rd cell (Cc: Benefit <br>4th cell (D): Concern <br>5th cell (E): Perception – Fight climate change <br>6th cell (F): Perception – Climate-friendly technology <br>7th cell (G): Perception – Extensive R&D needed <br>8th cell (H): Perception – Better options than CCS <br>9th cell (I): Sentiment <br>10th cell (J): Relatedness <br>11th cell (K): Comments </p> <p><strong>- Annotations </strong><br><strong>Benefit </strong><br>Preventing c. change <br>Reducing c. change risks <br>Safeguarding jobs <br>Creating new jobs <br>Fossil energy production envir. friendly <br>Products envir. friendly <br>Reducing envir. impact <br>Other <br>None <br><strong>Concern </strong><br>Accidents <br>Leakages <br>Environmental <br>Earthquake-related <br>Increased local traffic <br>Investment <br>Greenwashing <br>Lock-in effects for fossil energy <br>Increase cost <br>Other <br>None <br><strong>Perception (Yes / No / None) </strong><br>Fight climate change <br>Climate-friendly technology <br>Extensive R&D needed <br>Better options than CCS <br><strong>Sentiment </strong><br>Positive <br>Negative <br>Neutral </p>
Experimental datasets for sentiment analysis and emotion mining - Emotion Mining Toolkit (EMTk)
<p><strong>Description</strong></p> <p>Datasets for sentiment analysis and emotion mining, distributed with the Emotion Mining Toolkit (EMTk) Docker container (see <a href="https://collab-uniba.github.io/EMTk">https://collab-uniba.github.io/EMTk</a> for more):</p> <ul> <li>Stack Overflow - A couple of gold standards of 4,000+ posts, manually annotated for mining both emotions and polarity.</li> <li>Jira - A gold standard of ~4,000 issues, manually annotated for emotions.</li> </ul> <p><strong>Citation</strong></p> <p>Please, see the references below for the papers to cite. Do not cite this Zenodo upload directly.</p>
Sentiment analysis of tech media articles using VADER package and co-occurrence analysis
<p><strong>Sentiment analysis of tech media articles using VADER package and co-occurrence analysis</strong></p> <p><strong>Sources</strong>: Above 140k articles (01.2016-03.2019):</p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood's scores would be positive, but the negative term would bring the paragraph's score down.</p> <p>The presented tables include the most extreme co-occurring terms for the analysed social issue. The examples are chosen from the list of words with 30 most positive and 30 most negative sentiment. The presented graphs show the evolution of sentiments for social issues. The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p> <p> </p> <p><strong>Files</strong></p> <p>sentiments_mod11.csv sentiment score based on chosen unigrams</p> <p>sentiments_mod22.csv sentiment score based on chosen bigrams</p> <p>sentiments_cooc_mod11.csv, sentiments_cooc_mod12.csv, sentiments_cooc_mod21.csv, sentiments_cooc_mod22.csv combinations of co-occurrences: unigrams-unigrams, unigrams-bigrams, bigrams-unigrams, bigrams-bigrams</p> <p> </p>
AWARE: Dataset for Aspect-Based Sentiment Analysis of Apps Reviews
<p> </p> <p><em><strong>The peer-reviewed paper of AWARE dataset is published in ASEW 2021, and can be accessed through: <a href="http://doi.org/10.1109/ASEW52652.2021.00049">http://doi.org/10.1109/ASEW52652.2021.00049</a>. Kindly cite this paper when using AWARE dataset.</strong></em></p> <p> </p> <p>Aspect-Based Sentiment Analysis (ABSA) aims to identify the opinion (sentiment) with respect to a specific aspect. Since there is a lack of <em>smartphone apps reviews</em> dataset that is annotated to support the ABSA task, we present AWARE: <strong>A</strong>BSA <strong>W</strong>arehouse of <strong>A</strong>pps <strong>RE</strong>views.</p> <p>AWARE contains apps reviews from three different domains (Productivity, Social Networking, and Games), as each domain has its distinct functionalities and audience. Each sentence is annotated with three labels, as follows: </p> <ul> <li><strong>Aspect Term: </strong>a term that exists in the sentence and describes an aspect of the app that is expressed by the sentiment. A term value of “N/A” means that the term is not explicitly mentioned in the sentence.</li> <li><strong>Aspect Category:</strong> one of the pre-defined set of domain-specific categories that represent an aspect of the app (e.g., security, usability, etc.).</li> <li><strong>Sentiment:</strong> positive or negative.</li> </ul> <p><em>Note: games domain does not contain aspect terms.</em></p> <p>We provide a comprehensive dataset of 11323 sentences from the three domains, where each sentence is additionally annotated with a Boolean value indicating whether the sentence expresses a positive/negative opinion. In addition, we provide three separate datasets, one for each domain, containing only sentences that express opinions. The file named “AWARE_metadata.csv” contains a description of the dataset’s columns.</p> <p><strong>How AWARE can be used?</strong></p> <p>We designed AWARE such that it can be used to serve various tasks. The tasks can be, but are not limited to:</p> <ul> <li>Sentiment Analysis.</li> <li>Aspect Term Extraction.</li> <li>Aspect Category Classification.</li> <li>Aspect Sentiment Analysis.</li> <li>Explicit/Implicit Aspect Term Classification.</li> <li>Opinion/Not-Opinion Classification.</li> </ul> <p>Furthermore, researchers can experiment with and investigate the effects of different domains on users' feedback.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.