Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
141
datasets available to search
ShareScore release 0.9.0
Dataset results
141 results for “Sentiment”
Dataset : Advanced sentiment analysis using BERT
Open the record for dataset details and reuse information.
Data from: In the mood: the dynamics of collective sentiments on Twitter
Open the record for dataset details and reuse information.
Data from: Impact of lexical and sentiment factors on the popularity of scientific papers
Open the record for dataset details and reuse information.
Data from: Lexicon-enhanced sentiment analysis framework using rule-based classification scheme
Open the record for dataset details and reuse information.
Replication Package for "On Conclusion Validity of Empirical SE-studies with Sentiment Analysis Tools: Is Platform-Specific Retraining Enough?"
<p>This is the replication package for the paper "On Conclusion Validity of Empirical SE-studies with Sentiment Analysis Tools: Is Platform-Specific Retraining Enough?"</p>
Meet XLM-RLnews-8: Not Just Another Sentiment Analysis Model
Open the record for dataset details and reuse information.
Dataset: SWP-SentiSurvey for Sentiment Analysis in Student Software Projects
<p><strong>Description</strong></p> <p>In 2022, we conducted a survey about the perceptions of student developers regarding sentiments in statements. We published a paper about the results. The dataset includes the survey questions and the answers of the total 81 participants.</p> <p><strong>Citation</strong></p> <p>Information on the study design and execution are presented in the paper linked below.</p> <p>Please, see also the references below for the papers to cite.</p>
German consumer sentiment data from McMenamin et al., Institutions and Elections, Journalism, 2021.
<p>German consumer sentiment data (STATA) from McMenamin et al., Institutions and Elections, Journalism, 2021.</p> <p>Code and other datasets for this article also available on Zenodo.</p>
Sentiment analysis of media's political bias in micro-blogging - A computational framework for optimized recommendation systems
<p>This dataset contains tweets from four Pakistani news channels, namely DAWN, GEO, ARY, and 24News, collected using the twarc command line tool with a Twitter academic researcher account. The tweets were collected between Nov, 2015, and April 4, 2022, and relate to three major political parties in Pakistan, PTI, PMLN, and PPP through their official names in the relevant news channel page tweets. </p>
Volatility and Heterogeneity of Vaccine Sentiments Means Continuous Monitoring is Needed When Measuring Message Effectiveness
ClinicalTrials.gov study NCT05499299. IPD Sharing: YES. Countries: 1. Publications: 0.
Data from: Relative deprivation and relative wealth enhances anti-immigrant sentiments: the v-curve re-examined
Open the record for dataset details and reuse information.
Sentiment Analysis using Natural Language Processing implemented in Large Language Model
Open the record for dataset details and reuse information.
British Sentimental Novel Corpus (BSNC)
<p><span>Corpus de novela decimonónica inglesa —Romanticismo (1798–1836) y época victoriana (1837–1900)—<span> </span>formado por 114 novelas de 11 autores canónicos: Anthony Trollope, Charles Dickens, Charlotte Brontë, Elizabeth Gaskell, George Eliot, George Meredith, Jane Austen, Mary Shelley, Thomas Hardy, Walter Scott and William Makepeace Thackeray.</span></p> <p><span>Corpus of nineteenth-century British fiction —Romantic (1798–1836) and Victorian periods (1837–1900)— made up of 114 novels by 11 canonical authors: Anthony Trollope, Charles Dickens, Charlotte Brontë, Elizabeth Gaskell, George Eliot, George Meredith, Jane Austen, Mary Shelley, Thomas Hardy, Walter Scott and William Makepeace Thackeray.</span></p>
AckSent: Human Annotated Dataset of Support and Sentiments in Dissertation Acknowledgments
<p>This data is supplementary to the paper "<em>AckSent: Human Annotated Dataset of Support and Sentiments in Dissertation Acknowledgments" .<br></em></p>
Sentiment analysis in medication adherence: full code
<p>This database, associated with the research article on sentiment analysis in medication adherence, comprises an extensive collection of files essential for understanding and replicating the study's findings. It includes 362,806 anonymized medication reviews, meticulously analyzed to explore the correlation between patient sentiments and medication adherence.</p> <p><strong>Organize folders and code:</strong></p> <pre><code>MedicationAdherenceSentimentAnalysis/ │ ├── 01 - Original dataset/ │ └── originalDataset.csv │ ├── 02 - Dataset cleaning/ │ ├── CleaningStep1.py │ ├── CleaningStep2.py │ ├── CleaningStep3.py │ ├── CleaningStep4.py │ ├── Balancing.py │ ├── CleanedDatasetStep1.csv │ ├── CleanedDatasetStep2.csv │ ├── CleanedDatasetStep3.csv │ └── CleanedDatasetStep4.csv │ └── FinalBalancedDataset.csv │ ├── 03 - Vader analysis/ │ ├── vaderAnalysis.py │ └── vaderResults.csv │ ├── 04 - DistilRoBERTa analysis/ │ ├── DistilRoBERTa_Analysis.py │ └── DistilRoBERTa_Results.csv │ ├── 05 - Dataset report/ │ ├── datasetReport.py │ ├── filtredNegativeMeds.py │ ├── filtredPositiveMeds.py │ └── Outputs printed in terminal (Note: Terminal outputs are not stored as files) │ ├── 06 - Charts/ │ ├── charts.py │ ├── sentimentByLikert.html │ ├── vaderVsDistilRoBERTa.html │ ├── emotionsByLevelOfEffectiveness.html │ ├── emotionsByLevelOfEaseofuse.html │ └── emotionsByLevelOfSatisfaction.html │ └── 07 - Model metrics/ ├── modelMetrics.py └── dataserForMetrics.csv </code></pre> <p> </p> <p><strong>Contents:</strong></p> <ol> <li>Anonymized Medication Reviews: The original dataset must be downloaded from the original source indicated in "ReadmeDataset.docx". This is comprehensive dataset of 362,806 medication reviews, anonymized to protect patient privacy. These reviews serve as the primary data source for the sentiment analysis conducted in the study.</li> <li>Sentiment Analysis Results: Detailed files containing the output of the sentiment analysis performed using VADER and DistilRoBERTa models. This includes sentiment polarities, emotional responses, and their correlation with the perceived effectiveness, ease of use, and satisfaction reported by patients.</li> <li>Statistical Analysis Files: Contains the statistical tests and analysis results that establish the significant correlations between sentiment polarities and patient perceptions, as discussed in the article.</li> <li>Methodology Documentation: Detailed documentation of the methodologies used, including the application of AI tools like VADER and DistilRoBERTa for sentiment analysis. This section aids in replicating the study's approach for further research.</li> <li>Supplementary Information: Additional files that support the article's content, possibly including code snippets used for analysis, raw data processing details, and any other supplementary materials that contribute to the transparency and reproducibility of the research.</li> </ol> <p><strong>Usage:</strong></p> <p>This database is intended for researchers, clinicians, and academicians interested in exploring the intersection of artificial intelligence, sentiment analysis, and clinical pharmacy. It provides a rich resource for understanding patient sentiments towards medications and their potential impact on adherence. Researchers can utilize this data to replicate the study, conduct further analyses, or explore new hypotheses in the realm of health informatics and patient care optimization.</p>
Webis Cross-Lingual Sentiment Dataset 2010 (Webis-CLS-10)
<p>The Cross-Lingual Sentiment (CLS) dataset comprises about 800.000 Amazon product reviews in the four languages English, German, French, and Japanese.</p> <p>For more information on the construction of the dataset see (Prettenhofer and Stein, 2010) or the enclosed readme files. If you have a question after reading the paper and the readme files, please contact <a href="https://weimar.webis.de/people#prettenhofer">Peter Prettenhofer</a>.</p> <p>We provide the dataset in two formats: 1) a processed format which corresponds to the preprocessing (tokenization, etc.) in (Prettenhofer and Stein, 2010); 2) an unprocessed format which contains the full text of the reviews (e.g., for machine translation or feature engineering).</p> <p>The dataset was first used by (Prettenhofer and Stein, 2010). It consists of Amazon product reviews for three product categories---books, dvds and music---written in four different languages: English, German, French, and Japanese. The German, French, and Japanese reviews were crawled from Amazon in November, 2009. The English reviews were sampled from the <a href="http://www.cs.jhu.edu/~mdredze/datasets/sentiment/">Multi-Domain Sentiment Dataset</a> (Blitzer et. al., 2007). For each language-category pair there exist three sets of training documents, test documents, and unlabeled documents. The training and test sets comprise 2.000 documents each, whereas the number of unlabeled documents varies from 9.000 - 170.000.</p>
SEN - Sentiment analysis of Entities in News headlines
<p>If you wish to use this data please cite:</p> <p>Katarzyna Baraniak, Marcin Sydow,<br> A dataset for Sentiment analysis of Entities in News headlines (SEN),<br> Procedia Computer Science,<br> Volume 192,<br> 2021,<br> Pages 3627-3636,<br> ISSN 1877-0509,<br> <a href="https://doi.org/10.1016/j.procs.2021.09.136">https://doi.org/10.1016/j.procs.2021.09.136</a>.<br> (<a href="https://www.sciencedirect.com/science/article/pii/S1877050921018755">https://www.sciencedirect.com/science/article/pii/S1877050921018755</a>)</p> <p>bibtex: <a href="http://users.pja.edu.pl/~msyd/bibtex/sydow-baraniak-SENdataset-kes21.bib">users.pja.edu.pl/~msyd/bibtex/sydow-baraniak-SENdataset-kes21.bib</a><br> </p> <p>SEN is a novel publicly available human-labelled dataset for training and testing machine learning algorithms for the problem of entity level sentiment analysis of political news headlines.</p> <p>On-line news portals play a very important role in the information society. Fair media should present reliable and objective information. In practice there is an observable positive or negative bias concerning named entities (e.g. politicians) mentioned in the on-line news headlines.<br> Our dataset consists of 3819 human-labelled political news headlines coming from several major on-line media outlets in English and Polish.</p> <p>Each record contains a news headline, a named entity mentioned in the headline and a human annotated label (one of “positive”, “neutral”, “negative” ). Our SEN dataset package consists of 2 parts: SEN-en (English headlines that split into SEN-en-R and SEN-en-AMT), and SEN-pl (Polish headlines). Each headline-entity pair was annotated via team of volunteer researchers (the whole SEN-pl dataset and a subset of 1271 English records: the SEN-en-R subset, “R” for “researchers”) or via the Amazon Mechanical Turk service (a subset of 1360 English records: the SEN-en-AMT subset).</p> <p>During analysis of annotation outlying annotations and removed . Separate version of dataset without outliers is marked by "noutliers" in data file name.</p> <p>Details of the process of preparing the dataset and presenting its analysis are presented in the paper.<br> </p> <p>In case of any questions, please contact one of the authors. Email adresses are in the paper.</p>
Dat4API.ABSA: A Dataset of API Reviews from Stack Overflow for Aspect-Based Sentiment Analysis
<p>This is the dataset created from Stack Overflow discussions with manually labeled<br> Aspect-API-Sentiment information, used in the paper 'Dat4API.ABSA: A Dataset of API Reviews from Stack Overflow for Aspect-Based Sentiment Analysis'.</p>
MuSe-Wild: Multimodal Sentiment in-the-Wild Sub-challenge (MuSe2020)
<p><strong>MuSe-Wild of MuSe2020: </strong>Predicting the level of emotional dimensions (arousal, valence) in a time-continuous manner from audio-visual recordings. <strong>This package includes only MuSe-Wild features (all partitions) and annotations of the training and development set </strong>(test scoring via the MuSe website). </p> <p><strong>General: </strong>The purpose of the Multimodal Sentiment Analysis in Real-life media Challenge and Workshop (MuSe) is to bring together communities from different disciplines; mainly, the audio-visual emotion recognition community (signal-based), and the sentiment analysis community (symbol-based). </p> <p>We introduce the novel dataset MuSe-CAR that covers the range of aforementioned desiderata. MuSe-CAR is a large (>36h), multimodal dataset which has been gathered in-the-wild with the intention of further understanding Multimodal Sentiment Analysis in-the-wild, e.g., the emotional engagement that takes place during product reviews (i.e., automobile reviews) where a sentiment is linked to a topic or entity.</p> <p>We have designed MuSe-CAR to be of high voice and video quality, as informative video social media content, as well as everyday recording devices have improved in recent years. This enables robust learning, even with a high degree of novel, in-the-wild characteristics, for example as related to: i) Video: Shot size (a mix of close-up, medium, and long shots), face-angle (side, eye, low, high), camera motion (free, free but stable, and free but unstable, switch, e.g., zoom, fixed), reviewer visibility (full body, half-body, face only, and hands only), highly varying backgrounds, and people interacting with objects (car parts). ii) Audio: Ambient noises (car noises, music), narrator and host diarisation, diverse microphone types, and speaker locations. iii) Text: Colloquialisms, and domain-specific terms.</p> <p> </p>
Multilingual fine-grained sentiment analysis corpus
<p>A sentiment annotated corpus based on Fallout New Vegas. The corpus has the following sentiments: <em>neutral, anger, disgust, fear, happy, pained, sad, surprised</em> in the following languages: <em>English, German, Italian, Spanish and French</em>.</p> <p>Please cite the following paper: Mika Hämäläinen, Khalid Alnajjar, and Thierry Poibeau. 2022. Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog. In <em>FDG’22: Proceedings of the 17th International Conference on the Foundations of Digital Games (FDG ’22)</em></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.