Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

141

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

141 results for “Sentiment”

Learn how ShareScore rates datasets ↗
zenodo32/100

A COMPREHENSIVE STUDY OF MACHINE LEARNING APPROACHES FOR CUSTOMER SENTIMENT ANALYSIS IN BANKING SECTOR

<p>This study explores the application of sentiment analysis in the banking sector, focusing on customer feedback to enhance service quality and customer experiences. We collected a comprehensive dataset of approximately 100,000 entries from diverse sources, including customer satisfaction surveys, social media platforms, and direct feedback. A robust preprocessing pipeline was employed to address challenges associated with unstructured data, informal language, and mixed sentiments. We evaluated several machine learning and natural language processing models, including Logistic Regression, Naive Bayes, Support Vector Machine (SVM), Random Forest, Long Short-Term Memory (LSTM), and BERT (Bidirectional Encoder Representations from Transformers), using metrics such as accuracy, precision, recall, F1 score, AUC-ROC, and training time. The results revealed that advanced models, particularly BERT, achieved superior performance with an accuracy of 88% and an F1 score of 0.86, demonstrating an exceptional ability to capture nuanced sentiments. This study underscores the importance of employing sophisticated sentiment analysis techniques in banking to derive actionable insights from customer feedback. The findings suggest that leveraging advanced models can significantly improve service quality and customer satisfaction, while also presenting avenues for future research into real-time sentiment analysis and its integration with customer relationship management systems.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Dataset - Virtual Influencers and Public Perception: Social Media Sentiment Analysis and A Comprehensive Bibliometric

<p>The dataset contains a collection of articles related to Virtual Influencers and Public Perception: Social Media Sentiment Analysis and A Comprehensive Bibliometric in the Scopus database.</p>

opencc-by-4.0Nov 2024View details →
dryad32/100

Data from: Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types

<p>Concerns about gender bias in word embedding models have captured substantial attention in the algorithmic bias research literature. Other bias types however have received lesser amounts of scrutiny. This work describes a large-scale analysis of sentiment associations in popular word embedding models along the lines of gender and ethnicity but also along the less frequently studied dimensions of socioeconomic status, age, physical appearance, sexual orientation, religious sentiment and political leanings. Consistent with previous scholarly literature, this work has found systemic bias against given names popular among African-Americans in most embedding models examined. Gender bias in embedding models however appears to be multifaceted and often reversed in polarity to what has been regularly reported. Interestingly, using the common operationalization of the term bias in the fairness literature, novel types of so far unreported bias types in word embedding models have also been identified. Specifically, the popular embedding models analyzed here display negative biases against middle and working-class socioeconomic status, male children, senior citizens, plain physical appearance and intellectual phenomena such as Islamic religious faith, non-religiosity and conservative political orientation. Reasons for the paradoxical underreporting of these bias types in the relevant literature are probably manifold but widely held blind spots when searching for algorithmic bias and a lack of widespread technical jargon to unambiguously describe a variety of algorithmic associations could conceivably be playing a role. The causal origins for the multiplicity of loaded associations attached to distinct demographic groups within embedding models are often unclear but the heterogeneity of said associations and their potential multifactorial roots raises doubts about the validity of grouping them all under the umbrella term bias. Richer and more fine-grained terminology as well as a more comprehensive exploration of the bias landscape could help the fairness epistemic community to characterize and neutralize algorithmic discrimination more efficiently.</p>

opencc-zeroApr 2020View details →
zenodo32/100

DravidianMultiModality: A Dataset for Multi-modal Sentiment Analysis in Tamil and Malayalam

<p>@article{dravidian_multimodality,<br> &nbsp; title={DravidianMultiModality: A Dataset for Multi-modal Sentiment Analysis in Tamil and Malayalam},<br> &nbsp; author={Bharathi Raja Chakravarthi, Jishnu Parameswaran P.K, Premjith B, K.P Soman, Rahul Ponnusamy, Prasanna Kumar Kumaresan, Kingston Pal Thamburaj, John P. McCrae},<br> &nbsp; journal={arXiv.org},<br> &nbsp; publisher={2021}<br> }</p>

opencc-byJun 2021View details →
zenodo32/100

Data for manuscript: "Using Word Embeddings to Probe Sentiment Associations of Politically Loaded Terms in News and Opinion Articles from News Media Outlets"

<p>This data set contains material for the purpose of scientific reproducibility of the accompanying manuscript &quot;Using Word Embeddings to Probe Sentiment Associations of Politically Loaded Terms in News and Opinion Articles from News Media Outlets&quot;.</p> <p>Note that this data set is distributed with an Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) License. NonCommercial means you&nbsp;may not use the material for commercial purposes. NoDerivatives means if you remix, transform, or build upon the material, you may not distribute the modified material. Attribution means you must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use. See attached license terms for details.</p> <p>The work &quot;Using Word Embeddings to Probe Sentiment Associations of Politically Loaded Terms in News and Opinion Articles from News Media Outlets&quot; describes an analysis of political associations in 27 million diachronic (1975-2019) news and opinion articles from 47 news media outlets popular in the United States. We use embedding models trained on individual outlets content to quantify outlet-specific latent associations between positive/negative sentiment words and terms loaded with political connotations such as those describing political orientation, party affiliation, names of influential politicians and ideologically aligned public figures.&nbsp;</p> <p>News and opinion articles from the outlets listed in Figure 3 are available in the outlet&#39;s online domains and/or public cache repositories such as Google cache, The Internet Wayback Machine [31] and Common Crawl [32]. This work has not analyzed video or audio content of news media organizations, except when the outlet explicitly provides a transcript of such content in article form.<br> The temporal coverage of articles from different news outlets is not uniform. For most media organizations, news articles availability in their online domains or Internet cache backups becomes sparse as a function of articles&rsquo; age. This is not the case for some news outlets, where availability of news articles goes back to the 1970s. The Supplementary Material (SM) illustrates the time ranges of article data analyzed based on news outlets articles online availability.</p> <p>Textual content included in our analysis is circumscribed to the articles&rsquo; headlines and main text and does not include other article elements such as figure captions. Targeted textual content was located in HTML raw data using outlet specific XPath expressions. Tokens were lowercased prior to estimating embedding models. Markup language tags, URLs, nonalphanumeric characters, punctuation, digits, 330 common stop words and multiple spaces were removed prior to estimating word embeddings models.<br> All the analysis scripts and the diachronic word embedding models built from each of the 47 news media outlets analyzed in this work are available in this repository.</p> <p>For the purpose of reproducibility, we also provide in the above repository the articles&rsquo; text used to train the news outlets embedding models with the caveat that outlets articles not accessible without a subscription have been excluded. Also, for the included articles, stop words have been removed and the remaining words have been randomly scrambled within a sliding window of size 10 to render the articles incomprehensible to a human reader. These steps have been taken to not infringe articles copyright. These preprocessing steps have only minor impact on Continuous Bag of Words (CBOW) word2vec and the results reported in this work are similar when using the scrambled articles text to train outlet-specific embedding models.</p> <p>We derived outlet-specific word embedding models at every five-year time intervals within the 1975-2019 time range. The gensim [33] implementation of word2vec was used to train the embedding models. The continuous bag of words (CBOW) architecture performed slightly better than the Skip-Gram architecture in commonly used validation metrics so it was used for all subsequent analysis.&nbsp;</p> <p>For training the word embedding models, the following parameters were used: vector dimensions=300, window size=10, negative sampling=10, down sampling frequent words = 0.0001, minimum frequency count of 5 (only terms that appear more than 5 times in the corpus were included into the word embedding model vocabulary), number of training iterations (epochs) through the corpus=5. The exponent used to shape the negative sampling distribution was the default 0.75.&nbsp;</p> <p>Outlet-specific embedding models performance across a range of commonly used semantic, syntactic and analogy tasks was similar to popular pre-trained embedding models trained on corpora such as Twitter or Google books on similarity, association and word analogy tasks, see Supplemeentary Material of the manuscript for detailed validation tests results.</p> <p>&nbsp;</p>

opencc-by-nc-nd-4.0Jul 2021View details →
zenodo32/100

Video: Pandemic and Social Media Textual Sentiment Analysis of the Indonesian Goverment Policy in Facing the Thitd Wave of Covid-19 Attack

<p>This material has presented on 1st International Conference on Advance Research in Social and Economic Science. October 25, 2022</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Sentiment Inference: Pro and Contra relation dataset

<p>500 German sentences annotated for pro/con relations and polar roles of entities (negative/positive actors/effects): see References for a conceptual introduction.</p> <p>files: annotator1.conll .. annotator3.conll</p> <p>format: conll (parzu parser) with annotations</p> <p>- annotations at the end of the conll parse tree<br> &nbsp; - c = con<br> &nbsp; - p = pro<br> &nbsp; - neff,peff = negative, positive effect<br> &nbsp; - nac, pac = negative, positive actor<br> - the head indices are used for annotation (see below)<br> &nbsp; - c1,6 = Hofstetter con Gewerkschaften<br> &nbsp; - neff6 = negative Effekt on Gewerkschaften</p> <p><br> Note: in these annotations, pro/con is not an intentional relation</p> <p>- in &quot;Snow blocks the driveway&quot; it holds: con(snow,driveway)<br> - &quot;snow&quot; is a negative element wrt. to driveway<br> - use our animacy classifier to identify those case with an actor (see References lrec, available via IGGSA download)</p> <p><br> Example:<br> 1&nbsp;&nbsp; &nbsp;Hofstetter&nbsp;&nbsp; &nbsp;Hofstetter&nbsp;&nbsp; &nbsp;N&nbsp;&nbsp; &nbsp;NE&nbsp;&nbsp; &nbsp;_|Nom|Sg&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;subj&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 2&nbsp;&nbsp; &nbsp;wirft&nbsp;&nbsp; &nbsp;werfen&nbsp;&nbsp; &nbsp;V&nbsp;&nbsp; &nbsp;VVFIN&nbsp;&nbsp; &nbsp;3|Sg|Pres|Ind&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;root&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 3&nbsp;&nbsp; &nbsp;im&nbsp;&nbsp; &nbsp;in&nbsp;&nbsp; &nbsp;PREP&nbsp;&nbsp; &nbsp;APPRART&nbsp;&nbsp; &nbsp;Dat&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;pp&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 4&nbsp;&nbsp; &nbsp;Interview&nbsp;&nbsp; &nbsp;Interview&nbsp;&nbsp; &nbsp;N&nbsp;&nbsp; &nbsp;NN&nbsp;&nbsp; &nbsp;Neut|Dat|Sg&nbsp;&nbsp; &nbsp;3&nbsp;&nbsp; &nbsp;pn&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 5&nbsp;&nbsp; &nbsp;den&nbsp;&nbsp; &nbsp;die&nbsp;&nbsp; &nbsp;ART&nbsp;&nbsp; &nbsp;ART&nbsp;&nbsp; &nbsp;Def|Fem|Dat|Pl&nbsp;&nbsp; &nbsp;6&nbsp;&nbsp; &nbsp;det&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 6&nbsp;&nbsp; &nbsp;Gewerkschaften&nbsp;&nbsp; &nbsp;Gewerkschaft&nbsp;&nbsp; &nbsp;N&nbsp;&nbsp; &nbsp;NN&nbsp;&nbsp; &nbsp;Fem|Dat|Pl&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;objd&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 7&nbsp;&nbsp; &nbsp;vor&nbsp;&nbsp; &nbsp;vor&nbsp;&nbsp; &nbsp;PTKVZ&nbsp;&nbsp; &nbsp;PTKVZ&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;avz&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 8&nbsp;&nbsp; &nbsp;,&nbsp;&nbsp; &nbsp;,&nbsp;&nbsp; &nbsp;$,&nbsp;&nbsp; &nbsp;$,&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;root&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 9&nbsp;&nbsp; &nbsp;sie&nbsp;&nbsp; &nbsp;sie&nbsp;&nbsp; &nbsp;PRO&nbsp;&nbsp; &nbsp;PPER&nbsp;&nbsp; &nbsp;3|Pl|_|Nom&nbsp;&nbsp; &nbsp;10&nbsp;&nbsp; &nbsp;subj&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 10&nbsp;&nbsp; &nbsp;wollen&nbsp;&nbsp; &nbsp;wollen&nbsp;&nbsp; &nbsp;V&nbsp;&nbsp; &nbsp;VMFIN&nbsp;&nbsp; &nbsp;3|Pl|Pres|_&nbsp;&nbsp; &nbsp;2&nbsp;&nbsp; &nbsp;s&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 11&nbsp;&nbsp; &nbsp;die&nbsp;&nbsp; &nbsp;die&nbsp;&nbsp; &nbsp;ART&nbsp;&nbsp; &nbsp;ART&nbsp;&nbsp; &nbsp;Def|Fem|_|Sg&nbsp;&nbsp; &nbsp;12&nbsp;&nbsp; &nbsp;det&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 12&nbsp;&nbsp; &nbsp;Branche&nbsp;&nbsp; &nbsp;Branche&nbsp;&nbsp; &nbsp;N&nbsp;&nbsp; &nbsp;NN&nbsp;&nbsp; &nbsp;Fem|_|Sg&nbsp;&nbsp; &nbsp;13&nbsp;&nbsp; &nbsp;obja&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 13&nbsp;&nbsp; &nbsp;anschw&auml;rzen&nbsp;&nbsp; &nbsp;anschw&auml;rzen&nbsp;&nbsp; &nbsp;V&nbsp;&nbsp; &nbsp;VVINF&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;10&nbsp;&nbsp; &nbsp;aux&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> 14&nbsp;&nbsp; &nbsp;.&nbsp;&nbsp; &nbsp;.&nbsp;&nbsp; &nbsp;$.&nbsp;&nbsp; &nbsp;$.&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;0&nbsp;&nbsp; &nbsp;root&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;_&nbsp;&nbsp; &nbsp;<br> c1,6<br> p1,12<br> neff6</p> <p><br> References:</p> <p>@inproceedings{stance,<br> &nbsp; &nbsp; &nbsp; &nbsp;booktitle = {LSDSem 2017/LSD-Sem Linking Models of Lexical, Sentential and Discourse-level Semantics},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;month = {April},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;title = {Stance Detection in Facebook Posts of a German Right-wing Party},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; author = {Manfred Klenner and Don Tuggener and Simon Clematide},<br> &nbsp; &nbsp; &nbsp; &nbsp;publisher = {ResearchBib},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; year = {2017},<br> &nbsp; &nbsp; &nbsp; &nbsp; language = {english},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;url = {https://doi.org/10.5167/uzh-136567}<br> }<br> @inproceedings{perspectives,<br> &nbsp; &nbsp; &nbsp; &nbsp;booktitle = {18th International Conference on Computational Linguistics and Intelligent Text Processing},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;month = {April},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;title = {Verb-mediated Composition of Attitude Relations Comprising Reader and Writer Perspective},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; author = {Manfred Klenner and Simon Clematide and Don Tuggener},<br> &nbsp; &nbsp; &nbsp; &nbsp;publisher = {ResearchBib},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; year = {2017},<br> &nbsp; &nbsp; &nbsp; &nbsp; language = {english},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;url = {https://doi.org/10.5167/uzh-136569},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;doi = {10.1007/978-3-319-77116-8\_11}<br> }<br> @inproceedings{harmonization,<br> &nbsp; &nbsp; &nbsp; &nbsp;booktitle = {Proceedings of the 5th Swiss Text Analytics Conference (SwissText) \&amp; 16th Conference on Natural Language Processing (KONVENS)},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; editor = {Sarah Ebling and Don Tuggener and Manuela H{\&quot;u}rlimann and Martin Volk},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;month = {Juni 2020},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;title = {Harmonization Sometimes Harms},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; author = {Manfred Klenner and Anne G{\&quot;o}hring and Michael Amsler},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; publisher = {Virtual Event}<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; year = {2020},<br> &nbsp; &nbsp; &nbsp; &nbsp; language = {english},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;url = {https://doi.org/10.5167/uzh-197961}<br> }<br> @inproceedings{lrec,<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;month = {Juni},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; author = {Manfred Klenner and Anne G{\&quot;o}hring},<br> &nbsp; &nbsp; &nbsp; &nbsp;booktitle = {Proceedings of the Language Resources and Evaluation Conference},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;address = {Marseille, France},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;title = {Animacy Denoting {G}erman Nouns: Annotation and Classification},<br> &nbsp; &nbsp; &nbsp; &nbsp;publisher = {European Language Resources Association},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;pages = {1360--1364},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; year = {2022},<br> &nbsp; &nbsp; &nbsp; &nbsp; language = {english},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;url = {https://doi.org/10.5167/uzh-219148},<br> &nbsp; &nbsp; &nbsp; &nbsp; abstract = {In this paper, we introduce a gold standard for animacy detection comprising almost 14,500 German nouns that might be used to denote either animate entities or non-animate entities. We present inter-annotator agreement of our crowd-sourced seed annotations (9,000 nouns) and discuss the results of machine learning models applied to this data.}<br> }<br> &nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Sentiment Analysis for Arabic Tweets (Arabic)

<p>This video illustrates the idea of Opinion Mining and Sentiment Analysis for Arabic Tweets.&nbsp;&nbsp;Al Aisaee.F,&nbsp;Al Darmaki.Y and&nbsp;Al Rahbi.K undertook Twitter sentiment analysis with application in Arabic tweet text&nbsp;&nbsp;as a research project for their Bachelor final year project of Statistics major. The project was jointly supervised by Dr. Al Hasani.I and Dr. Zaidoom.H in 2019. The study focused on Arabic tweets about unemployment issue in Oman, using a trending hashtag on this topic at that time.&nbsp;</p> <p>Al Aisaee.F,&nbsp;Al Darmaki.Y and&nbsp;Al Rahbi.K designed this video to illustrate the concept of&nbsp;Opinion Mining and Sentiment Analysis for Arabic speakers.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
dryad32/100

Data from: Understanding sentiment of national park visitors from social media data

Open the record for dataset details and reuse information.

publicJul 2020View details →
dryad32/100

Data from: Wide range screening of algorithmic bias in word embedding models using large sentiment lexicons reveals underreported bias types

Open the record for dataset details and reuse information.

publicApr 2020View details →
zenodo28/100

News Title Sentiment Dataset

<p>This dataset is part of the Monash, UEA &amp;&nbsp;UCR time series regression repository.&nbsp;<a href="http://tseregression.org/">http://tseregression.org/</a></p> <p>The goal of this dataset is to predict sentiment score for news title.&nbsp;This dataset contains 83164 time series obtained from the News Popularity in Multiple Social Media Platforms dataset from the UCI repository.&nbsp;This is a large data set of news items and their respective social feedback on multiple platforms: Facebook, Google+ and LinkedIn.&nbsp;The collected data relates to a period of 8 months, between November 2015 and July 2016, accounting for about 100,000 news items on four different topics: economy, microsoft, obama and palestine.&nbsp;This data set is tailored for evaluative comparisons in predictive analytics tasks, although allowing for tasks in other research areas such as topic detection and tracking, sentiment analysis in short text, first story detection or news recommendation.&nbsp;The time series has 3 dimensions.&nbsp;<br> <br> Please refer to&nbsp;<a href="https://archive.ics.uci.edu/ml/datasets/News+Popularity+in+Multiple+Social+Media+Platforms">https://archive.ics.uci.edu/ml/datasets/News+Popularity+in+Multiple+Social+Media+Platforms</a>&nbsp;for more details<br> <br> Citation request<br> Nuno Moniz and Luis Torgo (2018), Multi-Source Social Feedback of Online News Feeds, CoRR</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

News Headline Sentiment Dataset

<p>This dataset is part of the Monash, UEA &amp;&nbsp;UCR time series regression repository.&nbsp;<a href="http://tseregression.org/">http://tseregression.org/</a></p> <p>The goal of this dataset is to predict sentiment score for news headline.&nbsp;This dataset contains 83164 time series obtained from the News Popularity in Multiple Social Media Platforms dataset from the UCI repository.&nbsp;This is a large data set of news items and their respective social feedback on multiple platforms: Facebook, Google+ and LinkedIn.&nbsp;The collected data relates to a period of 8 months, between November 2015 and July 2016, accounting for about 100,000 news items on four different topics: economy, microsoft, obama and palestine.&nbsp;This data set is tailored for evaluative comparisons in predictive analytics tasks, although allowing for tasks in other research areas such as topic detection and tracking, sentiment analysis in short text, first story detection or news recommendation.&nbsp;The time series has 3 dimensions.&nbsp;<br> <br> Please refer to <a href="https://archive.ics.uci.edu/ml/datasets/News+Popularity+in+Multiple+Social+Media+Platforms">https://archive.ics.uci.edu/ml/datasets/News+Popularity+in+Multiple+Social+Media+Platforms</a>&nbsp;for more details<br> <br> Citation request<br> Nuno Moniz and Luis Torgo (2018), Multi-Source Social Feedback of Online News Feeds, CoRR</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Sentiment Polarity of Programmers in an Open Source Software Project: An Exploratory Study

<p>Context: During the implementation of issues filed in open source software projects, programmers engage and interact in discussions on how to implement them. These discussions provide evidence to investigate emotionally loaded practices embraced by programmers. They interact to explain their point of view regarding the project and the issue under analysis. Objective: Analyze programmers sentiment polarity in an open source software project having releases as a reference for the analysis. Methods: We conducted an exploratory study to characterize the sentiment polarity of comments registered in issues associated with releases of the Moby open source software project. Results: The quantitative analysis identified sentiment polarity variations throughout consecutive releases in line with specific functionalities. Based on a qualitative analysis, we identified these functionalities and specific group of programmers that contributed to those results. Conclusions: We identified initial evidence to contribute for the understanding of the causes underlying the influence of the sentiments of the developers in the context of releases of open source software projects.</p>

opencc-by-4.0Oct 2020View details →
dryad28/100

Data from: Impact of lexical and sentiment factors on the popularity of scientific papers

We investigate how textual properties of scientific papers relate to the number of citations they receive. Our main finding is that correlations are nonlinear and affect differently the most cited and typical papers. For instance, we find that, in most journals, short titles correlate positively with citations only for the most cited papers, whereas for typical papers, the correlation is usually negative. Our analysis of six different factors, calculated both at the title and abstract level of 4.3 million papers in over 1500 journals, reveals the number of authors, and the length and complexity of the abstract, as having the strongest (positive) influence on the number of citations.

opencc-zeroDec 2015View details →
dryad28/100

Data from: In the mood: the dynamics of collective sentiments on Twitter

We study the relationship between the sentiment levels of Twitter users and the evolving network structure that the users created by @-mentioning each other. We use a large dataset of tweets to which we apply three sentiment scoring algorithms, including the open source SentiStrength program. Specifically we make three contributions. Firstly, we find that people who have potentially the largest communication reach (according to a dynamic centrality measure) use sentiment differently than the average user: for example, they use positive sentiment more often and negative sentiment less often. Secondly, we find that when we follow structurally stable Twitter communities over a period of months, their sentiment levels are also stable, and sudden changes in community sentiment from one day to the next can in most cases be traced to external events affecting the community. Thirdly, based on our findings, we create and calibrate a simple agent-based model that is capable of reproducing measures of emotive response comparable with those obtained from our empirical dataset.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Lexicon-enhanced sentiment analysis framework using rule-based classification scheme

With the rapid increase in social networks and blogs, the social media services are increasingly being used by online communities to share their views and experiences about a particular product, policy and event. Due to economic importance of these reviews, there is growing trend of writing user reviews to promote a product. Nowadays, users prefer online blogs and review sites to purchase products. Therefore, user reviews are considered as an important source of information in Sentiment Analysis (SA) applications for decision making. In this work, we exploit the wealth of user reviews, available through the online forums, to analyze the semantic orientation of words by categorizing them into +ive and -ive classes to identify and classify emoticons, modifiers, general-purpose and domain-specific words expressed in the public's feedback about the products. However, the un-supervised learning approach employed in previous studies is becoming less efficient due to data sparseness, low accuracy due to non-consideration of emoticons, modifiers, and presence of domain specific words, as they may result in inaccurate classification of users' reviews. Lexicon-enhanced sentiment analysis based on Rule-based classification scheme is an alternative approach for improving sentiment classification of users' reviews in online communities. In addition to the sentiment terms used in general purpose sentiment analysis, we integrate emoticons, modifiers and domain specific terms to analyze the reviews posted in online communities. To test the effectiveness of the proposed method, we considered users reviews in three domains. The results obtained from different experiments demonstrate that the proposed method overcomes limitations of previous methods and the performance of the sentiment analysis is improved after considering emoticons, modifiers, negations, and domain specific terms when compared to baseline methods.

opencc-zeroDec 2016View details →
zenodo28/100

SM-FEEL-BG Sentiments-Experiment 2 - No FB Dataset Splits

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo28/100

SENTIMENTAL PORTRAYAL IN EASTERN AND WESTERN LITERATURE

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Unveiling Developers' Feelings: A Benchmarking Study on Emotion and Sentiment Analysis of Software Commit Messages

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Understanding Media's Role in Public Perception: Discrepancies in Sentiment During the COVID-19 Pandemic in English and Croatian News Media

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record