Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
141
datasets available to search
ShareScore release 0.9.0
Dataset results
141 results for “Sentiment”
Produced Data of Naive Bayes Sentiment Classifier
<p>This is the data produced by the running of the Naive Bayes classifier algorithm. It is a list of every word in the vocabulary of the classifier, as well as the number of occurrences of each word, as well as the likelihood ratio of this word. Please note the likelihood ratio is calculated by taking the likelihood of word given a positive label divided by the likelihood of a word given a negative label. This data is licensed under the CC BY 4.0 international license, and may be taken and used freely with credit given. This data was produced by two different datasets, using a Naive Bayes classifier. These datasets were the Polarity Review v2.0 dataset from Cornell, and the Large Movie Review Dataset from Stanford.</p>
Data mining and sentiment analysis on Twitter and Facebook
<p>Les données récoltées sont sur le sujet "Data mining and sentiment analysis on Twitter and Facebook". Ce jeu de donnée contient la liste des attributs principaux suivants :</p> <ul> <li>titles, titre du fichier PDF,</li> <li>authors, auteurs du fichier PDF,</li> <li>years, année de création du fichier PDF,</li> <li>ncitedby, nombre de citation,</li> <li>linkfiles, liens du fichier PDF,</li> </ul> <p>mais également des métadonnées. </p> <p>La récupération du jeu de données a été récolté sur Google Scholar. Plusieurs recherches sur Google Scholar ont été faites pour ce dernier (voir liens ci-dessous) :</p> <ul> <li>https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=twitter+data+mining+filetype%3Apdf&btnG=</li> <li>https://scholar.google.com/scholar?start=490&q=facebook+data+mining+-Twitter+filetype:pdf&hl=en&as_sdt=0,5</li> <li>https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=seniment+analyse+twitter+filetype%3Apdf&btnG=</li> </ul>
Arabic news credibility on Twitter using sentiment analysis and ensemble learning
<p>Arabic news credibility on Twitter using sentiment analysis and ensemble learning.</p> <p> </p> <p>WHAT IS IT?</p> <p>-----------</p> <p>an Arabic news credibility model on Twitter using sentiment analysis and ensemble learning.</p> <p>Here we include the Collected dataset and the source code of the proposed model written in Python language and using Keras library with Tensorflow backend.</p> <p> </p> <p>Required Packages</p> <p>------------------</p> <ol> <li>Keras (<a href="https://keras.io/">https://keras.io/</a>).</li> <li>Scikit-learn (<a href="http://scikit-learn.org/)">http://scikit-learn.org/)</a></li> <li>Imnlearn (<a href="https://imbalanced-learn.org/stable/">imbalanced-learn documentation — Version 0.10.1</a>)</li> </ol> <p> </p> <p> </p> <p>To Run the model</p> <p>---------------</p> <p>One data file is required to run the model which are:</p> <p> </p> <ol> <li>The data that were used are the collected dataset in the file, set the path of the required data file in the code.</li> </ol> <p> </p> <p>The dataset</p> <p>---------------</p> <ol> <li>There are the dataset file with all features, you can choose the features that you need and apply it on the model.</li> <li>There are a description file that describe each feature in the news credibility dataset</li> <li>The file Tweet_ID contains the list of tweets id in the dataset.</li> <li>The annotated replies based on credibility is provided.</li> </ol> <p> </p> <p> </p> <p> </p> <p> </p> <p>CONTACTS</p> <p>--------</p> <ul> <li>If you want to report bugs or have general queries email to <duha_atif@yahoo.com></li> </ul> <p> </p> <p> </p> <p> </p>
Preprocessed Indonesian Twitter Dataset on UU Perlindungan Data Pribadi for Sentiment Analysis Research
<p>This dataset, titled 'Preprocessed Indonesian Twitter Dataset on UU Perlindungan Data Pribadi for Sentiment Analysis Research,' is curated and prepared for the purpose of conducting sentiment analysis research as outlined in the project 'ANALISIS SENTIMEN MASYARAKAT TERHADAP UU PERLINDUNGAN DATA PRIBADI PADA APLIKASI X DENGAN METODE SUPPORT VECTOR MACHINE' (Sentiment Analysis of the Community Towards the Personal Data Protection Law on Application X Using Support Vector Machine Method).</p>
Sentiment Quantification Datasets
<p>These files are contain the tokenized reviews that are used for quantication experiments on text.</p> <p>IMDB is derived from the IMDB dataset from Maas et al., 2011 (https://ai.stanford.edu/~amaas/data/sentiment/).<br> The version of the IMDB content in this dataset has minimal processing with respect to the original dataset, yet, it is provided to unsure reproducibility of experiments.</p> <p>HP and Kindle dataset are Amazon reviews collected by the authors. The reviews are respectively about the books in the Harry Potter series, and about the Kindle e-book reader.</p>
The Impacts of Sentiments and Tones in Community-Generated Issue Discussions
<p>The dataset, data analysis code, and complete results accompanying the paper published in CHASE2021, titled <em>The Impacts of Sentiments and Tones in Community-Generated Issue Discussions</em>.</p>
Hansard Speeches and Sentiment V1.0
<p>A public dataset of speeches in the Hansard, the record of the speeches, votes and legislation in the UK Parliament. The dataset provides information on each speech of ten words or longer, made in the House of Commons between 1980 and 2016, with information on the speaking MP, their party, gender and age at the time of the speech. The dataset also includes all speeches of ten words made from 1936 to 1979, without identifying information on the speaker.</p> <p>The speeches have been classified for sentiment using a total of five libraries from the R packages `sentimentr`, `syuzhet` and `lexicon`.</p> <p>The integrity of the public Hansard record is questionable at times, and while I have improved it, the data is presented 'as is'. More details on the dataset are available at: http://evanodell.com/datasets/hansard-data/</p>
Sentiment study of the digitalization - survey results
<p>Survey results on sentiment conducted in Poland.</p>
BDFoodSent: A Large-Scale Sentiment-Labeled Restaurant Review Dataset from Bangladesh
<p>BDFoodReview is a large-scale dataset containing 334,119 restaurant reviews collected from "Foodpanda Bangladesh". The dataset includes customer reviews in mixed languages (Bangla, English, and Banglish), translated into English, along with their corresponding ratings and sentiment labels.</p> <p> </p> <h3>Dataset Statistics</h3> <p>Total Reviews: 334,119</p> <p>Features/Columns: 19</p> <p> </p> <h3>Potential Applications</h3> <p>Sentiment Analysis</p> <p>Restaurant Review Classification</p> <p>Customer Satisfaction Analysis</p> <p>Opinion Mining</p> <p>Natural Language Processing Research</p> <p>Food Service Industry Analysis</p>
Data and R script for "Fear and cultural background drive sexual prejudice in France – A sentiment analysis approach"
<p>Data:</p> <p>corpus_integral.csv</p> <p>FEEL_1.csv</p> <p>mauvais.txt</p> <p>neg_hetero_corrected.txt</p> <p>participant_info_used.txt</p> <p>pos_hetero_corrected.txt</p> <p>R script:</p> <p>polarities.R</p> <p>sentiments_discrete.R</p>
Investor Sentiment and Information Efficiency: Evidence from the Bitcoin Market.
<p>This dataset contains information about the volume of Google searches for the term "bitcoin," the values of the Twitter happiness index and the Fear Index, the number of page views for the term "Bitcoin" on Wikipedia, and the online participation of the Bitcointalk.org online forum. Additionally, bitcoin prices and historical returns for the period July 2015 to June 2021 are included.</p>
Data for manuscript: "Longitudinal Analysis of Sentiment and Emotion in News Media Headlines Using Automated Labelling with Transformer Language Models"
<p>This data set contains automated sentiment and emotionality annotations of 23 million headlines from 47 popular news media outlets popular in the United States. </p> <p>The set of 47 news media outlets analysed (listed in Figure 1 of the main manuscript) was derived from the AllSides organization <a href="https://www.allsides.com/blog/updated-allsides-media-bias-chart-version-11">2019 Media Bias Chart v1.1</a>. The human ratings of outlets’ ideological leanings were also taken from this chart and are listed in Figure 2 of the main manuscript. </p> <p>News articles headlines from the set of outlets analyzed in the manuscript are available in the outlets’ online domains and/or public cache repositories such as The Internet Wayback Machine, Google cache and Common Crawl. Articles headlines were located in articles’ HTML raw data using outlet-specific XPath expressions. </p> <p>The temporal coverage of headlines across news outlets is not uniform. For some media organizations, news articles availability in online domains or Internet cache repositories becomes sparse for earlier years. Furthermore, some news outlets popular in 2019, such as <em>The Huffington Post</em> or <em>Breitbart</em>, did not exist in the early 2000’s. Hence, our data set is sparser in headlines sample size and representativeness for earlier years in the 2000-2019 timeline. Nevertheless, 18 outlets in our data set have chronologically continuous partial or full headline data availability fulfilling our inclusive criteria (see manuscript Methods) since the year 2000. Figure S 1 in the SI reports the number of headlines per outlet and per year in our analysis.</p> <p>In a small percentage of articles, outlet specific XPath expressions might fail to properly capture the content of the headline due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. After manual testing, we determined that the percentage of headlines following in this category is very small. Additionally, our method might miss detecting some articles in the online domains of news outlets. To conclude, in a data analysis of over 23 million headlines, we cannot manually check the correctness of every single data instance and hundred percent accuracy at capturing headlines’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our headlines set is representative of headlines in print news media content for the studied time period and outlets analyzed.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript as well as aggregated data of sentiment and emotionality automated annotations of the headlines and human annotations of a subset of headlines sentiment and emotionality used as ground truth. </p> <p>-models.rar contains the Transformer sentiment and emotion annotation models used in the analysis. Namely: </p> <p>Siebert/sentiment-roberta-large-english from https://huggingface.co/siebert/sentiment-roberta-large-english. This model is a fine-tuned checkpoint of <a href="https://huggingface.co/roberta-large">RoBERTa-large</a> (<a href="https://arxiv.org/pdf/1907.11692.pdf">Liu et al. 2019</a>). It enables reliable binary sentiment analysis for various types of English-language text. For each instance, it predicts either positive (1) or negative (0) sentiment. The model was fine-tuned and evaluated on 15 data sets from diverse text sources to enhance generalization across different types of texts (reviews, tweets, etc.). See more information from the original authors at https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>DistilbertSST2.rar is the default sentiment classification model of the HuggingFace Transformer library https://huggingface.co/ This model is only used to replicate the results of the sentiment analysis with sentiment-roberta-large-english </p> <p>DistilRoberta j-hartmann/emotion-english-distilroberta-base from https://huggingface.co/j-hartmann/emotion-english-distilroberta-base. The model is a fine-tuned checkpoint of <a href="https://huggingface.co/distilroberta-base">DistilRoBERTa-base</a>. The model allows annotation of English text with Ekman's 6 basic emotions, plus a neutral class. The model was trained on 6 diverse datasets. Please refer to the original author at https://huggingface.co/j-hartmann/emotion-english-distilroberta-base for an overview of the data sets used for fine tuning. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromSentimentRobertaLargeModel.rar URLs of headlines analyzed and the sentiment annotations of the siebert/sentiment-roberta-large-english Transformer model. https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromDistilbertSST2.rar URLs of headlines analyzed and the sentiment annotations of the default HuggingFace sentiment analysis model fine-tuned on the SST-2 dataset. https://huggingface.co/</p> <p>-headlinesDataWithEmotionLabelsAnnotationsFromDistilRoberta.rar URLs of headlines analyzed and the emotion categories annotations of the j-hartmann/emotion-english-distilroberta-base Transformer model. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p>
DMP Examine the correlation between Twitter Sentiment data and the stock data of the 4 big tech companies Apple, Amazon, Google and Microsoft
<p><span>The purpose of using these specific datasets are, the possibilities they give to analyze the stock data changes, based on the sentiment analysis of the previous day. Including this output information it is possible to analyze our goal of searching for possible correlation between those two. </span></p>
Annotated Dataset for Bilingual Code-Mixed English-Malay Sentiment Analysis and Sarcasm Detection in Public Security Domain
<p>Tweets from X, and post with comment from TikTok was acquired <span>from 11 September until 21 September 2022</span>. Data from both platforms was merged and selected. Three annotators manually labelling the selected data for sentiment and sarcasm. Sentiment labels are ‘positive’, ‘negative’, and ‘neutral’. Sarcasm label is ‘sarcastic’ and ‘not sarcastic’. Majority voting is considered for each label. Language identification label produced for each data. </p>
A Labelled Dataset for Sentiment Analysis of Videos on YouTube, TikTok, and other sources about the 2024 outbreak of Measles
<p><strong>Please cite the following paper when using this dataset:</strong></p> <p>N. Thakur, V. Su, M. Shao, K. Patel, H. Jeong, V. Knieling, and A. Bian “A labelled dataset for sentiment analysis of videos on YouTube, TikTok, and other sources about the 2024 outbreak of measles,” Proceedings of the 26th International Conference on Human-Computer Interaction (HCII 2024), Washington, USA, 29 June - 4 July 2024. (Accepted as a Late Breaking Paper, Preprint Available at: <a href="https://doi.org/10.48550/arXiv.2406.07693" rel="nofollow">https://doi.org/10.48550/arXiv.2406.07693</a>)</p> <p><strong>Abstract</strong></p> <p>This dataset contains the data of 4011 videos about the ongoing outbreak of measles published on 264 websites on the internet between January 1, 2024, and May 31, 2024. These websites primarily include YouTube and TikTok, which account for 48.6% and 15.2% of the videos, respectively. The remainder of the websites include Instagram and Facebook as well as the websites of various global and local news organizations. For each of these videos, the URL of the video, title of the post, description of the post, and the date of publication of the video are presented as separate attributes in the dataset. After developing this dataset, sentiment analysis (using VADER), subjectivity analysis (using TextBlob), and fine-grain sentiment analysis (using DistilRoBERTa-base) of the video titles and video descriptions were performed. This included classifying each video title and video description into (i) one of the sentiment classes i.e. positive, negative, or neutral, (ii) one of the subjectivity classes i.e. highly opinionated, neutral opinionated, or least opinionated, and (iii) one of the fine-grain sentiment classes i.e. fear, surprise, joy, sadness, anger, disgust, or neutral. These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for performing sentiment analysis or subjectivity analysis in this field as well as for other applications. The paper associated with this dataset (please see the above-mentioned citation) also presents a list of open research questions that may be investigated using this dataset.</p>
sentiment specific word embedding
<p>sentiment specific word embedding learned based on the approach described in the following paper:</p> <p>D. Tang, et al., Learning Sentiment-Specific Word Embedding for Twitter Sentiment Classification, ACL 2014</p> <p> </p>
Sentiment analysis of tech media articles using VADER package and co-occurrence analysis (01.2016-04.2019)
<p>Sentiment analysis of tech media articles using VADER package and co-occurrence analysis</p> <p><strong>Sources with weights:</strong></p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p><strong>Methodology</strong></p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood's scores would be positive, but the negative term would bring the paragraph's score down.</p> <p>The presented tables include the most extreme co-occurring terms for the analysed social issue. The examples are chosen from the list of words with 30 most positive and 30 most negative sentiment. The presented graphs show the evolution of sentiments for social issues. The analysed paragraphs are selected the following way:</p> <ul> <li>The articles containing the given social issue are identified</li> <li>The paragraphs containing the social issue are selected for sentiment analysis</li> </ul> <p>*Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>
Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis
<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, “Mpox narrative on Instagram: A labeled multilingual dataset of Instagram posts on mpox for sentiment, hate speech, and anxiety analysis,” arXiv [cs.LG], 2024, URL: https://arxiv.org/abs/2409.05292</p> <p><strong>Abstract</strong></p> <p>The world is currently experiencing an outbreak of mpox, which has been declared a Public Health Emergency of International Concern by WHO. During recent virus outbreaks, social media platforms have played a crucial role in keeping the global population informed and updated regarding various aspects of the outbreaks. As a result, in the last few years, researchers from different disciplines have focused on the development of social media datasets focusing on different virus outbreaks. No prior work in this field has focused on the development of a dataset of Instagram posts about the mpox outbreak. The work presented in this paper (stated above) aims to address this research gap. It presents this <strong>multilingual dataset of</strong> <strong>60,127 Instagram posts</strong> about mpox, published between <strong>July 23, 2022, and September 5, 2024</strong>. This dataset contains Instagram posts about mpox in <strong>52 languages</strong>. For each of these posts, the Post ID, Post Description, Date of publication, language, and translated version of the post (translation to English was performed using the Google Translate API) are presented as separate attributes in the dataset.</p> <p>After developing this dataset, sentiment analysis, hate speech detection, and anxiety or stress detection were also performed. This process included classifying each post into</p> <ul> <li>one of the fine-grain sentiment classes, i.e., <strong>fear, surprise, joy, sadness, anger, disgust, or neutral</strong>, </li> <li><strong>hate or not hate</strong></li> <li><strong>anxiety/stress detected or no anxiety/stress detected</strong>.</li> </ul> <p>These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for sentiment, hate speech, and anxiety or stress detection, as well as for other applications. </p> <p><strong>The 52 distinct languages in which Instagram posts are present in the dataset </strong><strong>are </strong>English, Portuguese, Indonesian, Spanish, Korean, French, Hindi, Finnish, Turkish, Italian, German, Tamil, Urdu, Thai, Arabic, Persian, Tagalog, Dutch, Catalan, Bengali, Marathi, Malayalam, Swahili, Afrikaans, Panjabi, Gujarati, Somali, Lithuanian, Norwegian, Estonian, Swedish, Telugu, Russian, Danish, Slovak, Japanese, Kannada, Polish, Vietnamese, Hebrew, Romanian, Nepali, Czech, Modern Greek, Albanian, Croatian, Slovenian, Bulgarian, Ukrainian, Welsh, Hungarian, and Latvian. </p> <p>The following table represents the data description for this dataset</p> <table> <tbody> <tr> <td> <p><strong>Attribute Name</strong></p> </td> <td> <p><strong>Attribute Description</strong></p> </td> </tr> <tr> <td> <p>Post ID</p> </td> <td> <p>Unique ID of each Instagram post</p> </td> </tr> <tr> <td> <p>Post Description</p> </td> <td> <p>Complete description of each post in the language in which it was originally published</p> </td> </tr> <tr> <td> <p>Date</p> </td> <td> <p>Date of publication in MM/DD/YYYY format</p> </td> </tr> <tr> <td> <p>Language</p> </td> <td> <p>Language of the post as detected using the Google Translate API</p> </td> </tr> <tr> <td> <p>Translated Post Description</p> </td> <td> <p>Translated version of the post description. All posts which were not in English were translated into English using the Google Translate API. No language translation was performed for English posts.</p> </td> </tr> <tr> <td> <p>Sentiment</p> </td> <td> <p>Results of sentiment analysis (using translated Post Description) where each post was classified into one of the sentiment classes: fear, surprise, joy, sadness, anger, disgust, and neutral</p> </td> </tr> <tr> <td> <p>Hate</p> </td> <td> <p>Results of hate speech detection (using translated Post Description) where each post was classified as hate or not hate</p> </td> </tr> <tr> <td> <p>Anxiety or Stress</p> </td> <td> <p>Results of anxiety or stress detection (using translated Post Description) where each post was classified as stress/anxiety detected or no stress/anxiety detected.</p> </td> </tr> </tbody> </table>
UK consumer sentiment data for McMenamin et al., Institutions and Elections, Journalism, 2021.
<p>UK consumer sentiment data (STATA) for McMenamin et al., Institutions and Elections, Journalism, 2021.</p> <p>Code and other datasets for this article also available on Zenodo.</p>
Forecasting with news sentiment: Evidence with UK newspapers
<p>These are datasets of economic sentiments derived from Uk newspapers using a dictionary and support vector machines. For more information on the application refer to : </p> <p>Rambaccussing, D. and Kwiatkowski, A., 2020. Forecasting with news sentiment: Evidence with UK newspapers. <em>International Journal of Forecasting</em>, <em>36</em>(4), pp.1501-1516.</p> <p>https://www.sciencedirect.com/science/article/pii/S0169207020300595</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.