Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
74
datasets available to search
ShareScore release 0.9.0
Dataset results
74 results for “Sentiment Analysis”
Sentiment Analysis of Kampus Merdeka Policy
A project analyzing public sentiment on Kampus Merdeka through data collection, text preprocessing, sentiment classification, and performance evaluation.
Sentiment analysis of tech media articles using VADER package and co-occurrence analysis (01.2016-04.2021)
<p>Sentiment analysis of tech media articles using VADER package and co-occurrence analysis</p> <p>Sources with weights:</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood's scores would be positive, but the negative term would bring the paragraph's score down.</p> <p> </p> <p>*Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>
Dataset: Sentiment Analysis annotation of News headlines covering the Olympic legacy of Rio 2016 and London 2012 published by the Brazilian and British online media
<p>Dataset of 464 news headlines with sentiment manually annotated by a domain expert using the labels positive, negative and neutral. Data contains URLs for news articles published between 2004-2020 by the British and Brazilian media in English and Brazilian Portuguese covering the Olympic legacies of London 2012 and Rio 2016. Articles were collected from the news outlets’ websites using Google search engine.</p> <p>News outlets:</p> <ul> <li>The Guardian</li> <li>Daily Mail</li> <li>Globo</li> <li>Estadao</li> </ul>
Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text
<p>Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text covering the Olympic legacy of Rio 2016 and London 2012. Data was searched via Google search engine. It is composed of sentiment labels assigned to 1271 news articles in total.</p> <p><strong>News outlets:</strong></p> <ul> <li>BBC</li> <li>Daily Mail</li> <li>The Telegraph</li> <li>The Guardian</li> <li>Globo</li> <li>Estadao</li> <li>Folha de S. Paulo</li> </ul> <p><strong>Events covered by the articles:</strong></p> <ul> <li>London 2012 Olympic legacy</li> <li>Rio 2016 Olympic legacy</li> </ul> <p>All classifiers were used in texts in English. Text originally published in Portuguese by the Brazilian media were automatically translated.</p> <p><strong>Sentiment classifiers used:</strong></p> <ul> <li>Vader</li> <li>BERT (Trained on Amazon data)</li> <li>BERT (Trained on twitter data - 140)</li> </ul> <p>Each document (spreadsheet - xlsx) refers to one outlet and one event (London 2012 or Rio 2016).</p> <p><strong>How were labels assigned to the texts?</strong></p> <p>These labels are a combination of the three sentiment classifiers listed above. If two of them agree with the same label, then this label would be considered as right. Otherwise, the label ‘other’ was assigned.</p> <p>For news article body text: the proportion of sentences of each sentiment type was used to assign labels to the whole article instead of averaging the sentence scores. For example, if the proportion of sentences with negative labels is greater than 50%, then the article is assigned a negative label.</p> <p><strong>The documents are composed of the following columns:</strong></p> <ul> <li>Rank: the position of the article on Google search ranking</li> <li>Date: date of article's publication (DD/MM/YYYY)</li> <li>Link: article's link</li> <li>Title: article's title</li> <li>Sentiment_Title: final sentiment for article headline</li> <li>Sentiment_Text: final sentiment for article's body text</li> </ul> <p><em>PS: Documents do not include articles' body text. </em></p> <p><strong>Sentiment is presented in labels as follows:</strong></p> <ul> <li>Pos: Positive</li> <li>Neg: Negative</li> <li>Neutral: Neutral</li> <li>other: inconclusive - if each of the 3 classifiers assigned a different label to the article, the label 'other' was used. Therefore, 'other' identifies contradictory results.</li> </ul> <p> </p>
Brazilian tweets classified for sentiment analysis
<p>Brazilian tweets classified for sentiment analysis</p>
Dataset: SentiSurvey for Sentiment Analysis in Software Projects
<p><strong>Description</strong></p> <p>In 2022, we conducted a survey about the perceptions of developers regarding sentiments in statements. We published a paper about the results. The dataset includes the survey questions and the answers of the total 180 participants.</p> <p><strong>Citation</strong></p> <p>Information on the study design and execution are presented in the paper linked below.</p> <p>Please, see also the references below for the papers to cite.</p>
A cooperative deep learning model for stock market prediction using deep autoencoder and sentiment analysis
<p>This data is used for Stock Market Prediction. </p>
Sentiment Analysis of RUU PDP with Naive Bayes, Support Vector Machine, and Random Forest Classification Algorithm
<p>Dataset from the results of data crawling via Twitter which discusses the Rancangan Undang Undang Pelindungan Data Pribadi to be used in the sentiment analysis process. The dataset is divided into several parts according to the process executed on RapidMiner.</p>
Dataset for Twitter Sentiment Analysis on Criminal Data Propagation using Naive Bayes Algorithm
<p>This study presents a dataset tailored for conducting sentiment analysis on Twitter regarding the propagation of criminal data. Leveraging the Naive Bayes algorithm, the dataset aims to facilitate research into public perceptions surrounding the dissemination of criminal data on social media platforms. Through a currated collection of tweets, researchers can explore the nuanced sentiments and attitudes expressed by users in response to this phenomenon.</p>
Brussel mobility Twitter sentiment analysis CSV Dataset
<p>SSH CENTRE (Social Sciences and Humanities for Climate, Energy aNd Transport Research Excellence) is a Horizon Europe project, engaging directly with stakeholders across research, policy, and business (including citizens) to strengthen social innovation, SSH-STEM collaboration, transdisciplinary policy advice, inclusive engagement, and SSH communities across Europe, accelerating the EU’s transition to carbon neutrality. <br>SSH CENTRE is based in a range of activities related to Open Science, inclusivity and diversity – especially with regards Southern and Eastern Europe and different career stages – including: development of novel SSH-STEM collaborations to facilitate the delivery of the EU Green Deal; SSH knowledge brokerage to support regions in transition; and the effective design of strategies for citizen engagement in EU R&I activities. Outputs include action-led agendas and building stakeholder synergies through regular Policy Insight events.<br>This is captured in a high-profile virtual SSH CENTRE generating and sharing best practice for SSH policy advice, overcoming fragmentation to accelerate the EU’s journey to a sustainable future.<br>The documents uploaded here are part of WP2 whereby novel, interdisciplinary teams were provided funding to undertake activities to develop a policy recommendation related to EU Green Deal policy. Each of these policy recommendations, and the activities that inform them, will be written-up as a chapter in an edited book collection. Three books will make up this edited collection - one on climate, one on energy and one on mobility. <br>As part of writing a chapter for the SSH CENTRE book on ‘Mobility’, we set out to analyse the sentiment of users on Twitter regarding shared and active mobility modes in Brussels. This involved us collecting tweets between 2017-2022. A tweet was collected if it contained a previously defined mobility keyword (for example: metro) and either the name of a (local) politician, a neighbourhood or municipality, or a (shared) mobility provider. The files attached to this Zenodo webpage is a csv files containing the tweets collected.”. </p>
BRAIN Journal-Sentiment Analysis on Embedded Systems Blended Courses-Figure 3. Sentiment analysis on extracted themes
<p>Figure 3 is presenting the sentiment analysis results from the point of view of the themes extracted from the corpus. The same preoccupation for the cost of the course is revealed, but this time the fact that MOOCs are free is appreciated. Students perceive that an integration of MOOCs into blended courses leads to a rapid information of the topics, such a feature receiving a high positive score of +3.46. The detailed explanations in this blended approach received a positive impact from the students with a total score of +2.70, but also the gained knowledge is among the most highly rated corpus themes. </p>
BRAIN Journal-Sentiment Analysis on Embedded Systems Blended Courses-Figure 2. Twitter sentiment analysis results
<p> In order to validate our results, the next step was to extract the sentiment analysis from Tweeter’s tweets (Figure 2) which are based on blending embedded systems-related courses. The obtained polarity is positive, so this results shows not only that students appreciated this in a positive manner, but also that the proposed technique for integrating MOOCs into embedded systems courses is a viable one. </p>
BRAIN Journal-Sentiment Analysis on Embedded Systems Blended Courses-Figure 1. Semantria result
<p>In Figure 1 the Semantria output is presented, having a positive polarity, with a score of 0.218. What is interesting to note here are the keywords extracted from students’ feedback. They noticed the integration of MOOCs in the Embedded Systems course as positive due to the fact that the new information is perceived as easier and the gained knowledge seems to be valuable. Students are affected by too many concepts and also by the idea of paying for the course.</p>
SENTIMENT ANALYSIS OF CUSTOMER FEEDBACK IN THE BANKING SECTOR: A COMPARATIVE STUDY OF MACHINE LEARNING MODELS
<p><span>This study investigates the application of sentiment analysis to customer feedback in the banking sector, utilizing natural language processing (NLP) techniques and machine learning models to classify customer sentiments into positive, neutral, and negative categories. Feedback was sourced from online platforms, including bank websites, social media, and third-party review sites. Data preprocessing steps, such as tokenization, stemming, and feature extraction using TF-IDF, were employed to prepare the text for analysis. Various machine learning algorithms, including Logistic Regression, Random Forest, Support Vector Machine (SVM), Long Short-Term Memory (LSTM), and Naïve Bayes, were implemented and evaluated using metrics such as accuracy, precision, recall, and F1-score. The results show that LSTM outperformed all models with a 91% accuracy, followed closely by SVM at 89%. These findings demonstrate the potential of advanced machine learning techniques in accurately classifying sentiments and provide valuable insights into customer satisfaction and areas for improvement within the banking sector. Future work aims to further optimize models for better classification of neutral feedback and explore more advanced deep learning models, such as BERT.</span></p>
Dataset Public Opinion of UAE and Sentiment Analysis Process
<p>This article explains the UAE public opinion towards Indonesia in the era of President Jokowi's administration. In the digital era, in IR studies, it must be understood that public opinion influences a country's foreign policy, including in the economic field. The research team crawled data from Twitter (X), a data set (raw data), to find out the UAE's public opinion. The raw data was then processed using SVM machine learning to find out the UAE public's positive, negative, and neutral levels regarding Indonesia. After that, the results of public opinion were compared with the UAE's and Indonesia's economic cooperation to determine whether there was a relationship between public opinion and the level of economic cooperation between the two countries. So, the data in this study are the UAE public opinion data collection, which involved tracking and analyzing public sentiment in various UAE-based media outlets and sentiment analysis processes.</p>
Sentiment Analysis and Cross-lingual Word Embeddings for Endangered Languages
<p>A sentiment analyzer and cross-lingual word embeddings for endangered languages (e.g., Erzya, Moksha, Skolt Sami, Komi-Zyrian).</p>
Towards Generalization of Machine Learning Models: An Arabic Sentiment Analysis Dataset
<p>This data set consists of approximately 1.64 Million Arabic tweets (shared by their IDs) posted from 2009 to 2020, and their corresponding sentiment using a three-point classification system of Positive, Negative and Neutral/Mixed. No specific locations and/or keywords were specified throughout the data collection to obtain variation in the dialects and topics represented within the dataset. It is important to note that any biases in the proposed dataset in relation to the dialects and/or topics discussed were unintentional.</p> <p><strong>Please use the following citation if you use this data in a paper:</strong></p> <blockquote> <p>Abdaljalil, S., Hassanein, S., Mubarak, H., & Abdelali, A. (2023). Towards Generalization of Machine Learning Models: A Case Study of Arabic Sentiment Analysis. <strong>Proceedings of the International AAAI Conference on Web and Social Media, 17(1), 971-980.</strong></p> </blockquote> <p> </p> <p> </p>
Data mining and sentiment analysis on Twitter and Facebook
<p>Les données récoltées sont sur le sujet "Data mining and sentiment analysis on Twitter and Facebook". Ce jeu de donnée contient la liste des attributs principaux suivants :</p> <ul> <li>titles, titre du fichier PDF,</li> <li>authors, auteurs du fichier PDF,</li> <li>years, année de création du fichier PDF,</li> <li>ncitedby, nombre de citation,</li> <li>linkfiles, liens du fichier PDF,</li> </ul> <p>mais également des métadonnées. </p> <p>La récupération du jeu de données a été récolté sur Google Scholar. Plusieurs recherches sur Google Scholar ont été faites pour ce dernier (voir liens ci-dessous) :</p> <ul> <li>https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=twitter+data+mining+filetype%3Apdf&btnG=</li> <li>https://scholar.google.com/scholar?start=490&q=facebook+data+mining+-Twitter+filetype:pdf&hl=en&as_sdt=0,5</li> <li>https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=seniment+analyse+twitter+filetype%3Apdf&btnG=</li> </ul>
Arabic news credibility on Twitter using sentiment analysis and ensemble learning
<p>Arabic news credibility on Twitter using sentiment analysis and ensemble learning.</p> <p> </p> <p>WHAT IS IT?</p> <p>-----------</p> <p>an Arabic news credibility model on Twitter using sentiment analysis and ensemble learning.</p> <p>Here we include the Collected dataset and the source code of the proposed model written in Python language and using Keras library with Tensorflow backend.</p> <p> </p> <p>Required Packages</p> <p>------------------</p> <ol> <li>Keras (<a href="https://keras.io/">https://keras.io/</a>).</li> <li>Scikit-learn (<a href="http://scikit-learn.org/)">http://scikit-learn.org/)</a></li> <li>Imnlearn (<a href="https://imbalanced-learn.org/stable/">imbalanced-learn documentation — Version 0.10.1</a>)</li> </ol> <p> </p> <p> </p> <p>To Run the model</p> <p>---------------</p> <p>One data file is required to run the model which are:</p> <p> </p> <ol> <li>The data that were used are the collected dataset in the file, set the path of the required data file in the code.</li> </ol> <p> </p> <p>The dataset</p> <p>---------------</p> <ol> <li>There are the dataset file with all features, you can choose the features that you need and apply it on the model.</li> <li>There are a description file that describe each feature in the news credibility dataset</li> <li>The file Tweet_ID contains the list of tweets id in the dataset.</li> <li>The annotated replies based on credibility is provided.</li> </ol> <p> </p> <p> </p> <p> </p> <p> </p> <p>CONTACTS</p> <p>--------</p> <ul> <li>If you want to report bugs or have general queries email to <duha_atif@yahoo.com></li> </ul> <p> </p> <p> </p> <p> </p>
Preprocessed Indonesian Twitter Dataset on UU Perlindungan Data Pribadi for Sentiment Analysis Research
<p>This dataset, titled 'Preprocessed Indonesian Twitter Dataset on UU Perlindungan Data Pribadi for Sentiment Analysis Research,' is curated and prepared for the purpose of conducting sentiment analysis research as outlined in the project 'ANALISIS SENTIMEN MASYARAKAT TERHADAP UU PERLINDUNGAN DATA PRIBADI PADA APLIKASI X DENGAN METODE SUPPORT VECTOR MACHINE' (Sentiment Analysis of the Community Towards the Personal Data Protection Law on Application X Using Support Vector Machine Method).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.