Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

74

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

74 results for “Sentiment Analysis”

Learn how ShareScore rates datasets ↗
zenodo36/100

Data and R script for "Fear and cultural background drive sexual prejudice in France – A sentiment analysis approach"

<p>Data:</p> <p>corpus_integral.csv</p> <p>FEEL_1.csv</p> <p>mauvais.txt</p> <p>neg_hetero_corrected.txt</p> <p>participant_info_used.txt</p> <p>pos_hetero_corrected.txt</p> <p>R script:</p> <p>polarities.R</p> <p>sentiments_discrete.R</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Data for manuscript: "Longitudinal Analysis of Sentiment and Emotion in News Media Headlines Using Automated Labelling with Transformer Language Models"

<p>This data set contains automated sentiment and emotionality annotations of 23 million headlines from 47 popular news media outlets popular in the United States.&nbsp;</p> <p>The set of 47 news media outlets analysed (listed in Figure 1&nbsp;of the main manuscript) was derived from the AllSides organization <a href="https://www.allsides.com/blog/updated-allsides-media-bias-chart-version-11">2019 Media Bias Chart v1.1</a>. The human ratings of outlets&rsquo; ideological leanings were also taken from this chart and are listed in Figure 2 of the main manuscript.&nbsp;</p> <p>News articles headlines from the set of outlets analyzed in the manuscript are available in the outlets&rsquo; online domains and/or public cache repositories such as The Internet Wayback Machine, Google cache and Common Crawl. Articles headlines were located in articles&rsquo; HTML raw data using outlet-specific XPath expressions.&nbsp;</p> <p>The temporal coverage of headlines across news outlets is not uniform. For some media organizations, news articles availability in online domains or Internet cache repositories becomes sparse for earlier years. Furthermore, some news outlets popular in 2019, such as <em>The Huffington Post</em> or <em>Breitbart</em>, did not exist in the early 2000&rsquo;s. Hence, our data set is sparser in headlines sample size and representativeness for earlier years in the 2000-2019 timeline. Nevertheless, 18 outlets in our data set have chronologically continuous partial or full headline data availability fulfilling our inclusive criteria (see manuscript Methods) since the year 2000.&nbsp;Figure S 1 in the SI&nbsp;reports the number of headlines per outlet and per year in our analysis.</p> <p>In a small percentage of articles, outlet specific XPath expressions might fail to properly capture the content of the headline due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. After manual testing, we determined that the percentage of headlines following in this category is very small.&nbsp;Additionally, our method might miss detecting some articles in the online domains of news outlets. To conclude, in a data analysis of over 23 million&nbsp;headlines, we cannot manually check the correctness of every single data instance and hundred percent accuracy at capturing headlines&rsquo; content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our headlines set is representative of headlines in print news media content for the studied time period and outlets analyzed.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript as well as aggregated data of sentiment and emotionality automated annotations of the headlines and human annotations of a subset of headlines sentiment and emotionality used as ground truth.&nbsp;</p> <p>-models.rar contains the Transformer sentiment and emotion annotation models used in the analysis. Namely:&nbsp;</p> <p>Siebert/sentiment-roberta-large-english from&nbsp;https://huggingface.co/siebert/sentiment-roberta-large-english.&nbsp;This model is a fine-tuned checkpoint of&nbsp;<a href="https://huggingface.co/roberta-large">RoBERTa-large</a>&nbsp;(<a href="https://arxiv.org/pdf/1907.11692.pdf">Liu et al. 2019</a>). It enables reliable binary sentiment analysis for various types of English-language text. For each instance, it predicts either positive (1) or negative (0) sentiment. The model was fine-tuned and evaluated on 15 data sets from diverse text sources to enhance generalization across different types of texts (reviews, tweets, etc.). See more information from the original authors at&nbsp;https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>DistilbertSST2.rar is the default sentiment classification model of the HuggingFace Transformer library&nbsp;https://huggingface.co/ This model is only used to replicate the results of the sentiment analysis with&nbsp;sentiment-roberta-large-english&nbsp;</p> <p>DistilRoberta&nbsp;j-hartmann/emotion-english-distilroberta-base from&nbsp;https://huggingface.co/j-hartmann/emotion-english-distilroberta-base. The model is a fine-tuned checkpoint of&nbsp;<a href="https://huggingface.co/distilroberta-base">DistilRoBERTa-base</a>. The model allows annotation of English text with&nbsp;&nbsp;Ekman&#39;s 6 basic emotions, plus a neutral class.&nbsp;The model was trained on 6 diverse datasets. Please refer to the original author at&nbsp;https://huggingface.co/j-hartmann/emotion-english-distilroberta-base for an overview of the data sets used for fine tuning.&nbsp;https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromSentimentRobertaLargeModel.rar URLs of headlines analyzed and the sentiment annotations of the&nbsp;siebert/sentiment-roberta-large-english Transformer model.&nbsp;https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromDistilbertSST2.rar&nbsp;URLs of headlines analyzed and the sentiment annotations of the default HuggingFace sentiment analysis model fine-tuned on the SST-2 dataset.&nbsp;https://huggingface.co/</p> <p>-headlinesDataWithEmotionLabelsAnnotationsFromDistilRoberta.rar URLs of headlines analyzed and the emotion categories annotations of the&nbsp;j-hartmann/emotion-english-distilroberta-base Transformer model.&nbsp;https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Annotated Dataset for Bilingual Code-Mixed English-Malay Sentiment Analysis and Sarcasm Detection in Public Security Domain

<p>Tweets from X, and post with comment from TikTok was acquired <span>from 11 September until 21 September 2022</span>. Data from both platforms was merged and selected. Three annotators manually labelling the selected data for sentiment and sarcasm. Sentiment labels are &lsquo;positive&rsquo;, &lsquo;negative&rsquo;, and &lsquo;neutral&rsquo;. Sarcasm label is &lsquo;sarcastic&rsquo; and &lsquo;not sarcastic&rsquo;. Majority voting is considered for each label. Language identification label produced for each data.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

A Labelled Dataset for Sentiment Analysis of Videos on YouTube, TikTok, and other sources about the 2024 outbreak of Measles

<p><strong>Please cite the following paper when using this dataset:</strong></p> <p>N. Thakur, V. Su, M. Shao, K. Patel, H. Jeong, V. Knieling, and A. Bian &ldquo;A labelled dataset for sentiment analysis of videos on YouTube, TikTok, and other sources about the 2024 outbreak of measles,&rdquo; Proceedings of the 26th International Conference on Human-Computer Interaction (HCII 2024), Washington, USA, 29 June - 4 July 2024. (Accepted as a Late Breaking Paper, Preprint Available at: <a href="https://doi.org/10.48550/arXiv.2406.07693" rel="nofollow">https://doi.org/10.48550/arXiv.2406.07693</a>)</p> <p><strong>Abstract</strong></p> <p>This dataset contains the data of 4011 videos about the ongoing outbreak of measles published on 264 websites on the internet between January 1, 2024, and May 31, 2024. These websites primarily include YouTube and TikTok, which account for 48.6% and 15.2% of the videos, respectively. The remainder of the websites include Instagram and Facebook as well as the websites of various global and local news organizations. For each of these videos, the URL of the video, title of the post, description of the post, and the date of publication of the video are presented as separate attributes in the dataset. After developing this dataset, sentiment analysis (using VADER), subjectivity analysis (using TextBlob), and fine-grain sentiment analysis (using DistilRoBERTa-base) of the video titles and video descriptions were performed. This included classifying each video title and video description into (i) one of the sentiment classes i.e. positive, negative, or neutral, (ii) one of the subjectivity classes i.e. highly opinionated, neutral opinionated, or least opinionated, and (iii) one of the fine-grain sentiment classes i.e. fear, surprise, joy, sadness, anger, disgust, or neutral. These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for performing sentiment analysis or subjectivity analysis in this field as well as for other applications. The paper associated with this dataset (please see the above-mentioned citation) also presents a list of open research questions that may be investigated using this dataset.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Sentiment analysis of tech media articles using VADER package and co-occurrence analysis (01.2016-04.2019)

<p>Sentiment analysis of tech media articles using VADER package and co-occurrence analysis</p> <p><strong>Sources with weights:</strong></p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5&nbsp;%</li> <li>IEEE Spectrum 5&nbsp;%</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood&#39;s scores would be positive, but the negative term would bring the paragraph&#39;s score down.</p> <p>The presented tables include the most extreme co-occurring terms for the analysed social issue. The examples are chosen from the list of words with 30 most positive and 30 most negative sentiment. The presented graphs show the evolution of sentiments for social issues. The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. &amp; Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis

<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, &ldquo;Mpox narrative on Instagram: A labeled multilingual dataset of Instagram posts on mpox for sentiment, hate speech, and anxiety analysis,&rdquo; arXiv [cs.LG], 2024, URL: https://arxiv.org/abs/2409.05292</p> <p><strong>Abstract</strong></p> <p>The world is currently experiencing an outbreak of mpox, which has been declared a Public Health Emergency of International Concern by WHO. During recent virus outbreaks, social media platforms have played a crucial role in keeping the global population informed and updated regarding various aspects of the outbreaks. As a result, in the last few years, researchers from different disciplines have focused on the development of social media datasets focusing on different virus outbreaks. No prior work in this field has focused on the development of a dataset of Instagram posts about the mpox outbreak. The work presented in this paper (stated above) aims to address this research gap. It presents this <strong>multilingual dataset of</strong>&nbsp;<strong>60,127 Instagram posts</strong> about mpox, published between <strong>July 23, 2022, and September 5, 2024</strong>. This dataset contains Instagram posts about mpox in <strong>52 languages</strong>. For each of these posts, the Post ID, Post Description, Date of publication, language, and translated version of the post (translation to English was performed using the Google Translate API) are presented as separate attributes in the dataset.</p> <p>After developing this dataset, sentiment analysis, hate speech detection, and anxiety or stress detection were also performed. This process included classifying each post into</p> <ul> <li>one of the fine-grain sentiment classes, i.e., <strong>fear, surprise, joy, sadness, anger, disgust, or neutral</strong>,&nbsp;</li> <li><strong>hate or not hate</strong></li> <li><strong>anxiety/stress detected or no anxiety/stress detected</strong>.</li> </ul> <p>These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for sentiment, hate speech, and anxiety or stress detection, as well as for other applications.&nbsp;</p> <p><strong>The 52 distinct languages in which Instagram posts are present in the dataset&nbsp;</strong><strong>are&nbsp;</strong>English, Portuguese, Indonesian, Spanish, Korean, French, Hindi, Finnish, Turkish, Italian, German, Tamil, Urdu, Thai, Arabic, Persian, Tagalog, Dutch, Catalan, Bengali, Marathi, Malayalam, Swahili, Afrikaans, Panjabi, Gujarati, Somali, Lithuanian, Norwegian, Estonian, Swedish, Telugu, Russian, Danish, Slovak, Japanese, Kannada, Polish, Vietnamese, Hebrew, Romanian, Nepali, Czech, Modern Greek, Albanian, Croatian, Slovenian, Bulgarian, Ukrainian, Welsh, Hungarian, and Latvian.&nbsp;</p> <p>The following table represents the data description for this dataset</p> <table> <tbody> <tr> <td> <p><strong>Attribute Name</strong></p> </td> <td> <p><strong>Attribute Description</strong></p> </td> </tr> <tr> <td> <p>Post ID</p> </td> <td> <p>Unique ID of each Instagram post</p> </td> </tr> <tr> <td> <p>Post Description</p> </td> <td> <p>Complete description of each post in the language in which it was originally published</p> </td> </tr> <tr> <td> <p>Date</p> </td> <td> <p>Date of publication in MM/DD/YYYY format</p> </td> </tr> <tr> <td> <p>Language</p> </td> <td> <p>Language of the post as detected using the Google Translate API</p> </td> </tr> <tr> <td> <p>Translated Post Description</p> </td> <td> <p>Translated version of the post description. All posts which were not in English were translated into English using the Google Translate API. No language translation was performed for English posts.</p> </td> </tr> <tr> <td> <p>Sentiment</p> </td> <td> <p>Results of sentiment analysis (using translated Post Description) where each post was classified into one of the sentiment classes: fear, surprise, joy, sadness, anger, disgust, and neutral</p> </td> </tr> <tr> <td> <p>Hate</p> </td> <td> <p>Results of hate speech detection (using translated Post Description) where each post was classified as hate or not hate</p> </td> </tr> <tr> <td> <p>Anxiety or Stress</p> </td> <td> <p>Results of anxiety or stress detection (using translated Post Description) where each post was classified as stress/anxiety detected or no stress/anxiety detected.</p> </td> </tr> </tbody> </table>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Sentiment analysis data and word embeddings for Erzya, Komi-Zyrian, Moksha and Udmurt

<p>The aligned sentiment annotated data is in setiment_eval_data.json, vectors.zip has the word embeddings in a textual Gensim format, code.zip has the code and models.zip the sentiment analysis model.</p> <p>Please cite the following paper:</p> <p><strong>Alnajjar, K., H&auml;m&auml;l&auml;inen, M., &amp; Rueter, J, (2023)&nbsp;Sentiment Analysis Using Aligned Word Embeddings for Uralic Languages. In <em>Proceedings of the Second Workshop on Resources and Representations for Under-resourced Languages and Domains (RESOURCEFUL-2023)</em></strong></p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Forex News Annotated Dataset for Sentiment Analysis

<p>This dataset contains&nbsp;news headlines relevant to key forex pairs: AUDUSD, EURCHF, EURUSD, GBPUSD, and USDJPY. The data was extracted from reputable platforms <a href="https://www.forexlive.com">Forex Live</a>&nbsp;and <a href="https://www.fxstreet.com/">FXstreet</a>&nbsp;over a period of 86 days, from January to May 2023. The&nbsp;dataset comprises 2,291 unique news headlines. Each headline includes an associated forex pair, timestamp, source, author, URL, and the corresponding article text. Data was collected using web scraping techniques executed via a custom service on a virtual machine. This service periodically retrieves the latest news for a specified forex pair (ticker) from each platform, parsing all available information. The collected data is then processed to extract details such as the article&#39;s timestamp, author, and URL. The URL is further used to retrieve the full text of each article. This data acquisition process repeats approximately every 15 minutes.</p> <p>To ensure the reliability of the dataset, we manually annotated each headline for sentiment. Instead of solely focusing on the textual content, we <strong>ascertained sentiment based on the potential short-term impact of the headline on its corresponding forex pair</strong>. This method recognizes the currency market&#39;s acute sensitivity to economic news, which significantly influences many trading strategies. As such, this dataset could serve as an invaluable resource for fine-tuning sentiment analysis models in the financial realm.</p> <p>We used three categories for annotation: &#39;positive&#39;, &#39;negative&#39;, and &#39;neutral&#39;, which correspond to bullish, bearish, and hold sentiments, respectively, for the forex pair linked to each headline. The following&nbsp;Table&nbsp;provides examples of annotated headlines along with brief explanations of the assigned sentiment.&nbsp;</p> Examples of Annotated Headlines Forex Pair Headline Sentiment Explanation GBPUSD&nbsp; Diminishing bets for a move to 12400&nbsp; Neutral Lack of strong sentiment in either direction GBPUSD&nbsp; No reasons to dislike Cable in the very near term as long as the Dollar momentum remains soft&nbsp;&nbsp; Positive Positive sentiment towards GBPUSD (Cable) in the near term GBPUSD&nbsp; When are the UK jobs and how could they affect GBPUSD &nbsp; Neutral Poses a question and does not express a clear sentiment JPYUSD Appropriate to continue monetary easing to achieve 2% inflation target with wage growth&nbsp;&nbsp; Positive Monetary easing from Bank of Japan (BoJ) could lead to a weaker JPY in the short term due to increased money supply USDJPY Dollar rebounds despite US data. Yen gains amid lower yields &nbsp; Neutral Since both the USD and JPY are gaining, the effects on the USDJPY forex pair might offset each other USDJPY USDJPY to reach 124 by Q4 as the likelihood of a BoJ policy shift should accelerate Yen gains &nbsp; Negative USDJPY is expected to reach a lower value, with the USD losing value against the JPY AUDUSD <p>RBA Governor Lowe&rsquo;s Testimony High inflation is damaging and corrosive &nbsp;</p> Positive Reserve Bank of Australia (RBA) expresses concerns about inflation. Typically, central banks combat high inflation with higher interest rates, which could strengthen AUD. <p>Moreover, the dataset includes two columns with the predicted sentiment class and score as predicted by the <a href="https://huggingface.co/ProsusAI/finbert">FinBERT</a> model. Specifically, the FinBERT model outputs a set of probabilities for each sentiment class (positive, negative, and neutral), representing the model&#39;s confidence in associating the input headline with each sentiment category. These probabilities are used to determine the predicted class and a sentiment score&nbsp;for each headline. The sentiment score is computed by subtracting the negative class probability from the positive one.</p>

opencc-by-4.0May 2023View details →
zenodo32/100

Dataset and scripts for "Sentiment Analysis over Collaborative Relationships in Open Source Software Projects"

<p>Dataset and scripts for &quot;Sentiment Analysis over Collaborative Relationships in Open Source Software Projects&quot;.</p> <p>README is included in the files</p>

opencc-by-4.0Jul 2019View details →
zenodo32/100

SemEval-2020 Task 9: Overview of Sentiment Analysis of Code-Mixed Tweets

<p>There are 2 sub-tasks: sentiment analysis for Spanglish (Spanish-English) and for Hinglish (Hindi-English).</p> <p>The sentiment classes are Positive, negative, neutral.&nbsp;</p> <p>Hinglish dataset has 20k instances.</p> <p>Spanglish dataset has ~19k instances.&nbsp;</p> <p>Website:&nbsp;<a href="https://ritual-uh.github.io/sentimix2020/">https://ritual-uh.github.io/sentimix2020/</a></p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

Aspect-based Sentiment Analysis of Scientific Reviews - Openreview dataset

<p>The dataset contains all the data used in the JCDL 2020 research paper: <a href="https://dl.acm.org/doi/10.1145/3383583.3398541">Aspect-based Sentiment Analysis of Scientific Reviews</a></p> <p>The dataset is split into multiple files containing&nbsp;all the sentence annotations and the ICLR open review dataset (with reviews and scores and the confidence scores, final recommendation, etc.) for the last three years.</p> <p>The file &quot;iclr_conf.p&quot; is a pickle file which contains a NumPy array object.<br> The array contains 2681 rows corresponding to each accepted or rejected paper of 2017,2018,2019<br> Each row contains 4 columns.<br> The first column is the link of the paper in openreview.net, from where the data related to the paper is collected.<br> The second column is either 0 or 1, corresponding to the final decision: rejection or acceptance respectively.<br> The third column is the year of the conference for the particular submission.<br> The fourth column is another NumPy array containing 3 reviews in 3 rows. Each row of this array contains 3 columns containing the list of sentences in the same sequence as it appears in the text of the review, the confidence(ranging from 1-5), and the rating(ranging(1-10)) respectively.</p> <p>Each line of the file &quot;sentences.csv&quot; contains one sentence whose corresponding annotation is provided in the corresponding line in the file &quot;annotations.csv&quot;<br> The file &quot;annotations.csv&quot; is a file containing 8 comma-separated integers in each line.<br> Each column corresponds to the following aspects: Appropriateness, Clarity, Originality, Empirical/Theoretical Soundness, Meaningful Comparison, Substance,<br> Impact of Dataset/Software/Ideas and Recommendation.<br> An integer 0,1,2,3 corresponds to the following sentiment labels of the sentence on that aspect: Absent, Positive, Negative, Neutral</p> <p>Please cite our paper published in JCDL-2020 if you use our data: <a href="https://dl.acm.org/doi/10.1145/3383583.3398541">https://dl.acm.org/doi/10.1145/3383583.3398541</a></p>

opencc-by-4.0Oct 2020View details →
zenodo32/100

Datasets of the article "From Classification to Quantification in Tweet Sentiment Analysis"

<p>Datasets used for the following SNAM paper:<br> ---------------------------------------------------------------------------------------------------<br> Title: From Classification to Quantification in Tweet Sentiment Analysis<br> Authors: Wei Gao and Fabrizio Sebastiani<br> Organization: Qatar Computing Research Institute, Hamad Bin Khalifa University, Doha, Qatar<br> ---------------------------------------------------------------------------------------------------</p> <p>[Content]</p> <p>* SemEval2013, SemEval2014, SemEval2015 datasets:<br> &nbsp; - semeval.train.feature.txt: Training set for learning sentiment models at development stage<br> &nbsp; - semeval.dev.feature.txt: Held-out set for tuning parameters<br> &nbsp; - semeval.train+dev.feature.txt: Training set for learning the final sentiment model<br> &nbsp; - semeval13.test.feature.txt: SemEval2013 test set<br> &nbsp; - semeval14.test.feature.txt: SemEval2014 test set<br> &nbsp; - semeval15.test.feature.txt: SemEval2015 test set<br> &nbsp;&nbsp;<br> * Other datasets: semeval2016, sanders, sst, omd, hcr, gasp, wa, wb<br> &nbsp; - X.train.feature.txt: Training set for learning sentiment models at development stage<br> &nbsp; - X.dev.feature.txt: Held-out set for tuning parameters<br> &nbsp; - X.train+dev.feature.txt: Training set for learning the final sentiment model<br> &nbsp; - X.test.feature.txt (or X.dev-test.feature.txt for semeval2016 only): Test set<br> where X is one of semeval2016, sanders, sst, omd, hcr and gasp.</p> <p>* Training files are saved in ./data/train directory, and held-out and test files are in ./data/test directory</p> <p><br> For more details, please refer to the paper.</p> <p><br> [Citation]<br> You can cite the following paper when referring to the dataset:</p> <pre>@article{gao2016classification, title={From classification to quantification in tweet sentiment analysis}, author={Gao, Wei and Sebastiani, Fabrizio}, journal={Social Network Analysis and Mining}, volume={6}, number={1}, pages={19}, year={2016}, publisher={Springer} }</pre> <p>&nbsp;</p>

opencc-by-4.0Nov 2015View details →
zenodo32/100

URDU Dataset for Multi-modal Sentiment Analysis

<p>The "Multi-modal Sentiment Analysis Dataset for Urdu Language Opinion Videos" is a valuable resource aimed at advancing research in sentiment analysis, natural language processing, and multimedia content understanding. This dataset is specifically curated to cater to the unique context of Urdu language opinion videos, a dynamic and influential content category in the digital landscape.</p> <p><strong>Dataset Description:</strong></p> <ul> <li><strong>Size and Diversity:</strong>&nbsp;This dataset comprises an extensive collection of Urdu language opinion videos, encompassing a wide spectrum of topics and sentiments. It consists of a total of 214 videos, each of varying lengths, offering a diverse and comprehensive representation of the Urdu language content landscape.</li> <li><strong>Sentiment Annotations:</strong>&nbsp;The dataset is meticulously annotated with sentiment labels, providing information on the emotional tone expressed in each video. The sentiment labels include "positive," "negative," and "neutral," offering a nuanced understanding of the sentiment conveyed in these multimedia opinion pieces.</li> <li><strong>Multi-modal Approach:</strong>&nbsp;A unique feature of this dataset is its multi-modal approach. It combines text, audio, and visual data to enable researchers to delve into the various dimensions of sentiment analysis within the context of opinion videos. The multi-modal annotations encompass the textual content of spoken words, the auditory characteristics of the videos, and the visual cues from the video frames.</li> </ul> <p><strong>Significance and Applications:</strong></p> <p>This dataset holds significant value for both the research community and practical applications:</p> <ul> <li><strong>Research Advancement:</strong>&nbsp;Researchers can employ this dataset to investigate the complex landscape of sentiment analysis within the context of opinion videos. It facilitates inquiries into sentiment trends, the development of sentiment analysis models, and the creation of sentiment-aware multimedia content analysis tools.</li> <li><strong>Content Recommendation:</strong> The dataset can play a pivotal role in the development of content recommendation systems that cater to viewers' emotional preferences. Understanding sentiment in opinion videos is crucial for improving content engagement and user experience.</li> <li><strong>User Engagement Analysis:</strong> The dataset can empower studies on user engagement and interaction with multimedia content. It is an essential resource for researchers aiming to decode the factors influencing viewer reactions and engagement in multimedia.</li> </ul> <p>&nbsp;</p> <p>Researchers are encouraged to explore and utilize this dataset for various academic and commercial purposes, fostering innovation in sentiment analysis and multimedia understanding. The dataset is made available with open access to facilitate collaborative research and to contribute to the broader knowledge in the field.</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

Dataset for Sentiment Analysis of X platform about "MK Hasil Pemilu"

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis

<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, &ldquo;Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis&rdquo;, Proceedings of the 7th International Conference on Machine Learning and Natural Language Processing (MLNLP 2024), Chengdu, China, October 18-20, 2024 (Paper accepted for publication, Preprint available at: https://arxiv.org/abs/2410.03293)</p> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p>The outbreak of COVID-19 served as a catalyst for content creation and dissemination on social media platforms, as such platforms serve as virtual communities where people can connect and communicate with one another seamlessly. While there have been several works related to the mining and analysis of COVID-19-related posts on social media platforms such as Twitter (or X), YouTube, Facebook, and TikTok, there is still limited research that focuses on the public discourse on Instagram in this context. Furthermore, the prior works in this field have only focused on the development and analysis of datasets of Instagram posts published during the first few months of the outbreak. The work presented in this paper aims to address this research gap and presents a novel multilingual dataset of <strong>500,153 Instagram posts about COVID-19 published between January 2020 and September 2024</strong>. This dataset contains Instagram posts in <strong>161 different languages</strong>. After the development of this dataset, multilingual sentiment analysis was performed using VADER and twitter-xlm-roberta-base-sentiment. This process involved classifying each post as positive, negative, or neutral. The results of sentiment analysis are presented as a separate attribute in this dataset.</p> <p><em><strong>For each of these posts, the Post ID, Post Description, Date of publication, language code, full version of the language, and sentiment label are presented as separate attributes in the dataset.</strong></em></p> <p>The Instagram posts in this dataset are present in <strong>161 different languages</strong> out of which the top 10 languages in terms of frequency are English (343041 posts), Spanish (30220 posts), Hindi (15832 posts), Portuguese (15779 posts), Indonesian (11491 posts), Tamil (9592 posts), Arabic (9416 posts), German (7822 posts), Italian (5162 posts), Turkish (4632 posts)</p> <p>There are <strong>535,021 distinct hashtags in this dataset</strong> with the top 10 hashtags in terms of frequency being #covid19 (169865 posts), #covid (132485 posts), #coronavirus (117518 posts), #covid_19 (104069 posts), #covidtesting (95095 posts), #coronavirusupdates (75439 posts), #corona (39416 posts), #healthcare (38975 posts), #staysafe (36740 posts), #coronavirusoutbreak (34567 posts)</p> <p>The following is a description of the attributes present in this dataset</p> <ul> <li><em><strong>Post ID</strong></em>:&nbsp;Unique ID of each Instagram post</li> <li><em><strong>Post Description</strong></em>:&nbsp;Complete description of each post in the language in which it was originally published</li> <li><em><strong>Date</strong></em>: Date of publication in MM/DD/YYYY format</li> <li><em><strong>Language code</strong></em>: Language code (for example: &ldquo;en&rdquo;) that represents the language of the post as detected using the Google Translate API&nbsp;</li> <li><em><strong>Full Language</strong></em>: Full form of the language (for example: &ldquo;English&rdquo;) that represents the language of the post as detected using the Google Translate API&nbsp;</li> <li><em><strong>Sentiment</strong></em>: Results of sentiment analysis (using the preprocessed version of each post) where each post was classified as positive, negative, or neutral</li> </ul> <p><strong>Open Research Questions</strong></p> <p>This dataset is expected to be helpful for the investigation of the following research questions and even beyond:</p> <ol> <li>How does sentiment toward COVID-19 vary across different languages?</li> <li>How has public sentiment toward COVID-19 evolved from 2020 to the present?</li> <li>How do cultural differences affect social media discourse about COVID-19 across various languages?</li> <li>How has COVID-19 impacted mental health, as reflected in social media posts across different languages?</li> <li>How effective were public health campaigns in shifting public sentiment in different languages?</li> <li>What patterns of vaccine hesitancy or support are present in different languages?</li> <li>How did geopolitical events influence public sentiment about COVID-19 in multilingual social media discourse?</li> <li>What role does social media discourse play in shaping public behavior toward COVID-19 in different linguistic communities?</li> <li>How does the sentiment of minority or underrepresented languages compare to that of major world languages regarding COVID-19?</li> <li>What insights can be gained by comparing the sentiment of COVID-19 posts in widely spoken languages (e.g., English, Spanish) to those in less common languages?</li> </ol> <p>All the Instagram posts that were collected during this data mining process to develop this dataset were publicly available on Instagram and did not require a user to log in to Instagram to view the same (at the time of writing this paper).</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

A COMPREHENSIVE STUDY OF MACHINE LEARNING APPROACHES FOR CUSTOMER SENTIMENT ANALYSIS IN BANKING SECTOR

<p>This study explores the application of sentiment analysis in the banking sector, focusing on customer feedback to enhance service quality and customer experiences. We collected a comprehensive dataset of approximately 100,000 entries from diverse sources, including customer satisfaction surveys, social media platforms, and direct feedback. A robust preprocessing pipeline was employed to address challenges associated with unstructured data, informal language, and mixed sentiments. We evaluated several machine learning and natural language processing models, including Logistic Regression, Naive Bayes, Support Vector Machine (SVM), Random Forest, Long Short-Term Memory (LSTM), and BERT (Bidirectional Encoder Representations from Transformers), using metrics such as accuracy, precision, recall, F1 score, AUC-ROC, and training time. The results revealed that advanced models, particularly BERT, achieved superior performance with an accuracy of 88% and an F1 score of 0.86, demonstrating an exceptional ability to capture nuanced sentiments. This study underscores the importance of employing sophisticated sentiment analysis techniques in banking to derive actionable insights from customer feedback. The findings suggest that leveraging advanced models can significantly improve service quality and customer satisfaction, while also presenting avenues for future research into real-time sentiment analysis and its integration with customer relationship management systems.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Dataset - Virtual Influencers and Public Perception: Social Media Sentiment Analysis and A Comprehensive Bibliometric

<p>The dataset contains a collection of articles related to Virtual Influencers and Public Perception: Social Media Sentiment Analysis and A Comprehensive Bibliometric in the Scopus database.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

DravidianMultiModality: A Dataset for Multi-modal Sentiment Analysis in Tamil and Malayalam

<p>@article{dravidian_multimodality,<br> &nbsp; title={DravidianMultiModality: A Dataset for Multi-modal Sentiment Analysis in Tamil and Malayalam},<br> &nbsp; author={Bharathi Raja Chakravarthi, Jishnu Parameswaran P.K, Premjith B, K.P Soman, Rahul Ponnusamy, Prasanna Kumar Kumaresan, Kingston Pal Thamburaj, John P. McCrae},<br> &nbsp; journal={arXiv.org},<br> &nbsp; publisher={2021}<br> }</p>

opencc-byJun 2021View details →
zenodo32/100

Video: Pandemic and Social Media Textual Sentiment Analysis of the Indonesian Goverment Policy in Facing the Thitd Wave of Covid-19 Attack

<p>This material has presented on 1st International Conference on Advance Research in Social and Economic Science. October 25, 2022</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Sentiment Analysis for Arabic Tweets (Arabic)

<p>This video illustrates the idea of Opinion Mining and Sentiment Analysis for Arabic Tweets.&nbsp;&nbsp;Al Aisaee.F,&nbsp;Al Darmaki.Y and&nbsp;Al Rahbi.K undertook Twitter sentiment analysis with application in Arabic tweet text&nbsp;&nbsp;as a research project for their Bachelor final year project of Statistics major. The project was jointly supervised by Dr. Al Hasani.I and Dr. Zaidoom.H in 2019. The study focused on Arabic tweets about unemployment issue in Oman, using a trending hashtag on this topic at that time.&nbsp;</p> <p>Al Aisaee.F,&nbsp;Al Darmaki.Y and&nbsp;Al Rahbi.K designed this video to illustrate the concept of&nbsp;Opinion Mining and Sentiment Analysis for Arabic speakers.&nbsp;</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record