Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
74
datasets available to search
ShareScore release 0.9.0
Dataset results
74 results for “Sentiment Analysis”
Data from: Lexicon-enhanced sentiment analysis framework using rule-based classification scheme
With the rapid increase in social networks and blogs, the social media services are increasingly being used by online communities to share their views and experiences about a particular product, policy and event. Due to economic importance of these reviews, there is growing trend of writing user reviews to promote a product. Nowadays, users prefer online blogs and review sites to purchase products. Therefore, user reviews are considered as an important source of information in Sentiment Analysis (SA) applications for decision making. In this work, we exploit the wealth of user reviews, available through the online forums, to analyze the semantic orientation of words by categorizing them into +ive and -ive classes to identify and classify emoticons, modifiers, general-purpose and domain-specific words expressed in the public's feedback about the products. However, the un-supervised learning approach employed in previous studies is becoming less efficient due to data sparseness, low accuracy due to non-consideration of emoticons, modifiers, and presence of domain specific words, as they may result in inaccurate classification of users' reviews. Lexicon-enhanced sentiment analysis based on Rule-based classification scheme is an alternative approach for improving sentiment classification of users' reviews in online communities. In addition to the sentiment terms used in general purpose sentiment analysis, we integrate emoticons, modifiers and domain specific terms to analyze the reviews posted in online communities. To test the effectiveness of the proposed method, we considered users reviews in three domains. The results obtained from different experiments demonstrate that the proposed method overcomes limitations of previous methods and the performance of the sentiment analysis is improved after considering emoticons, modifiers, negations, and domain specific terms when compared to baseline methods.
Unveiling Developers' Feelings: A Benchmarking Study on Emotion and Sentiment Analysis of Software Commit Messages
Open the record for dataset details and reuse information.
Dataset : Advanced sentiment analysis using BERT
Open the record for dataset details and reuse information.
Data from: Lexicon-enhanced sentiment analysis framework using rule-based classification scheme
Open the record for dataset details and reuse information.
Replication Package for "On Conclusion Validity of Empirical SE-studies with Sentiment Analysis Tools: Is Platform-Specific Retraining Enough?"
<p>This is the replication package for the paper "On Conclusion Validity of Empirical SE-studies with Sentiment Analysis Tools: Is Platform-Specific Retraining Enough?"</p>
Meet XLM-RLnews-8: Not Just Another Sentiment Analysis Model
Open the record for dataset details and reuse information.
Dataset: SWP-SentiSurvey for Sentiment Analysis in Student Software Projects
<p><strong>Description</strong></p> <p>In 2022, we conducted a survey about the perceptions of student developers regarding sentiments in statements. We published a paper about the results. The dataset includes the survey questions and the answers of the total 81 participants.</p> <p><strong>Citation</strong></p> <p>Information on the study design and execution are presented in the paper linked below.</p> <p>Please, see also the references below for the papers to cite.</p>
Sentiment analysis of media's political bias in micro-blogging - A computational framework for optimized recommendation systems
<p>This dataset contains tweets from four Pakistani news channels, namely DAWN, GEO, ARY, and 24News, collected using the twarc command line tool with a Twitter academic researcher account. The tweets were collected between Nov, 2015, and April 4, 2022, and relate to three major political parties in Pakistan, PTI, PMLN, and PPP through their official names in the relevant news channel page tweets. </p>
Sentiment Analysis using Natural Language Processing implemented in Large Language Model
Open the record for dataset details and reuse information.
Sentiment analysis in medication adherence: full code
<p>This database, associated with the research article on sentiment analysis in medication adherence, comprises an extensive collection of files essential for understanding and replicating the study's findings. It includes 362,806 anonymized medication reviews, meticulously analyzed to explore the correlation between patient sentiments and medication adherence.</p> <p><strong>Organize folders and code:</strong></p> <pre><code>MedicationAdherenceSentimentAnalysis/ │ ├── 01 - Original dataset/ │ └── originalDataset.csv │ ├── 02 - Dataset cleaning/ │ ├── CleaningStep1.py │ ├── CleaningStep2.py │ ├── CleaningStep3.py │ ├── CleaningStep4.py │ ├── Balancing.py │ ├── CleanedDatasetStep1.csv │ ├── CleanedDatasetStep2.csv │ ├── CleanedDatasetStep3.csv │ └── CleanedDatasetStep4.csv │ └── FinalBalancedDataset.csv │ ├── 03 - Vader analysis/ │ ├── vaderAnalysis.py │ └── vaderResults.csv │ ├── 04 - DistilRoBERTa analysis/ │ ├── DistilRoBERTa_Analysis.py │ └── DistilRoBERTa_Results.csv │ ├── 05 - Dataset report/ │ ├── datasetReport.py │ ├── filtredNegativeMeds.py │ ├── filtredPositiveMeds.py │ └── Outputs printed in terminal (Note: Terminal outputs are not stored as files) │ ├── 06 - Charts/ │ ├── charts.py │ ├── sentimentByLikert.html │ ├── vaderVsDistilRoBERTa.html │ ├── emotionsByLevelOfEffectiveness.html │ ├── emotionsByLevelOfEaseofuse.html │ └── emotionsByLevelOfSatisfaction.html │ └── 07 - Model metrics/ ├── modelMetrics.py └── dataserForMetrics.csv </code></pre> <p> </p> <p><strong>Contents:</strong></p> <ol> <li>Anonymized Medication Reviews: The original dataset must be downloaded from the original source indicated in "ReadmeDataset.docx". This is comprehensive dataset of 362,806 medication reviews, anonymized to protect patient privacy. These reviews serve as the primary data source for the sentiment analysis conducted in the study.</li> <li>Sentiment Analysis Results: Detailed files containing the output of the sentiment analysis performed using VADER and DistilRoBERTa models. This includes sentiment polarities, emotional responses, and their correlation with the perceived effectiveness, ease of use, and satisfaction reported by patients.</li> <li>Statistical Analysis Files: Contains the statistical tests and analysis results that establish the significant correlations between sentiment polarities and patient perceptions, as discussed in the article.</li> <li>Methodology Documentation: Detailed documentation of the methodologies used, including the application of AI tools like VADER and DistilRoBERTa for sentiment analysis. This section aids in replicating the study's approach for further research.</li> <li>Supplementary Information: Additional files that support the article's content, possibly including code snippets used for analysis, raw data processing details, and any other supplementary materials that contribute to the transparency and reproducibility of the research.</li> </ol> <p><strong>Usage:</strong></p> <p>This database is intended for researchers, clinicians, and academicians interested in exploring the intersection of artificial intelligence, sentiment analysis, and clinical pharmacy. It provides a rich resource for understanding patient sentiments towards medications and their potential impact on adherence. Researchers can utilize this data to replicate the study, conduct further analyses, or explore new hypotheses in the realm of health informatics and patient care optimization.</p>
SEN - Sentiment analysis of Entities in News headlines
<p>If you wish to use this data please cite:</p> <p>Katarzyna Baraniak, Marcin Sydow,<br> A dataset for Sentiment analysis of Entities in News headlines (SEN),<br> Procedia Computer Science,<br> Volume 192,<br> 2021,<br> Pages 3627-3636,<br> ISSN 1877-0509,<br> <a href="https://doi.org/10.1016/j.procs.2021.09.136">https://doi.org/10.1016/j.procs.2021.09.136</a>.<br> (<a href="https://www.sciencedirect.com/science/article/pii/S1877050921018755">https://www.sciencedirect.com/science/article/pii/S1877050921018755</a>)</p> <p>bibtex: <a href="http://users.pja.edu.pl/~msyd/bibtex/sydow-baraniak-SENdataset-kes21.bib">users.pja.edu.pl/~msyd/bibtex/sydow-baraniak-SENdataset-kes21.bib</a><br> </p> <p>SEN is a novel publicly available human-labelled dataset for training and testing machine learning algorithms for the problem of entity level sentiment analysis of political news headlines.</p> <p>On-line news portals play a very important role in the information society. Fair media should present reliable and objective information. In practice there is an observable positive or negative bias concerning named entities (e.g. politicians) mentioned in the on-line news headlines.<br> Our dataset consists of 3819 human-labelled political news headlines coming from several major on-line media outlets in English and Polish.</p> <p>Each record contains a news headline, a named entity mentioned in the headline and a human annotated label (one of “positive”, “neutral”, “negative” ). Our SEN dataset package consists of 2 parts: SEN-en (English headlines that split into SEN-en-R and SEN-en-AMT), and SEN-pl (Polish headlines). Each headline-entity pair was annotated via team of volunteer researchers (the whole SEN-pl dataset and a subset of 1271 English records: the SEN-en-R subset, “R” for “researchers”) or via the Amazon Mechanical Turk service (a subset of 1360 English records: the SEN-en-AMT subset).</p> <p>During analysis of annotation outlying annotations and removed . Separate version of dataset without outliers is marked by "noutliers" in data file name.</p> <p>Details of the process of preparing the dataset and presenting its analysis are presented in the paper.<br> </p> <p>In case of any questions, please contact one of the authors. Email adresses are in the paper.</p>
Dat4API.ABSA: A Dataset of API Reviews from Stack Overflow for Aspect-Based Sentiment Analysis
<p>This is the dataset created from Stack Overflow discussions with manually labeled<br> Aspect-API-Sentiment information, used in the paper 'Dat4API.ABSA: A Dataset of API Reviews from Stack Overflow for Aspect-Based Sentiment Analysis'.</p>
Multilingual fine-grained sentiment analysis corpus
<p>A sentiment annotated corpus based on Fallout New Vegas. The corpus has the following sentiments: <em>neutral, anger, disgust, fear, happy, pained, sad, surprised</em> in the following languages: <em>English, German, Italian, Spanish and French</em>.</p> <p>Please cite the following paper: Mika Hämäläinen, Khalid Alnajjar, and Thierry Poibeau. 2022. Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog. In <em>FDG’22: Proceedings of the 17th International Conference on the Foundations of Digital Games (FDG ’22)</em></p> <p> </p>
Unveiling the Sentiments and Opinions of Football Fans towards Video Assistant Referee (VAR) Technology: A Natural Language Processing Analysis
<p>The dataset has been used for the titled "Unveiling the Sentiments and Opinions of Football Fans towards Video Assistant Referee (VAR) Technology: A Natural Language Processing Analysis" provides valuable insights into the attitudes and opinions of football fans towards VAR technology. The dataset is the result of a natural language processing analysis of social media posts and online discussions related to VAR technology. It includes a comprehensive collection of sentiments, opinions, and attitudes expressed by football fans towards VAR technology, and provides researchers and analysts with a wealth of data to help them understand the public perception of this emerging technology in football. It can be used for further analysis in text analytics. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.