Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “Twitter, vaccination”
VaxxHesitancy: A Dataset for Studying Hesitancy Towards COVID-19 Vaccination on Twitter
<p>We create a publicly available dataset of over 3,100 COVID-19 vaccine-related tweets labeled as one of four stance categories: <em>pro-vaxx, anti-vaxx</em>, <em>vaxx-hesitant</em>,<em> or irrelevant</em>.</p> <p><strong>***</strong></p> <p><strong>Please use the V2 version.</strong></p> <p><strong>***</strong></p> <p>We split our dataset into two separate files:</p> <p>(1) VaccineHesitancy_train_v2.csv (Single + Double annotated)</p> <p>(2) VaccineHesitancy_test.csv (Double annotated)</p> <p>We present the details of this dataset here:</p> <p>VaxxHesitancy: A Dataset for Studying Hesitancy Towards COVID-19 Vaccination on Twitter (ICWSM 2023)</p> <p><strong>Our Pre-trained model</strong> (GateNLP/covid-vaccine-twitter-bert) : https://huggingface.co/GateNLP/covid-vaccine-twitter-bert</p> <p><strong>Paper</strong>: https://ojs.aaai.org/index.php/ICWSM/article/view/22213/21992</p> <p> </p> <pre>@inproceedings{mu2023vaxxhesitancy, title={VaxxHesitancy: A Dataset for Studying Hesitancy Towards COVID-19 Vaccination on Twitter}, author={Mu, Yida and Jin, Mali and Grimshaw, Charlie and Scarton, Carolina and Bontcheva, Kalina and Song, Xingyi}, booktitle={Proceedings of the International AAAI Conference on Web and Social Media}, volume={17}, pages={1052--1062}, year={2023} } </pre> <p> </p> <p> </p> <p> </p>
MAVIS Twitter dataset: A collection of tweets and sentiment analysis in Spanish about vaccines and diseases during the period 2015-2018
<p>MAVIS dataset comprises a full knowledge base regarding Twitter messages published in Spanish during the period 2015-2018, in the context of sentiment analysis of specific vaccines and their related diseases. Such diseases and vaccines are summarized as follows:</p> <ul> <li>Invasive meningococcal disease (“EMI” in Spanish): Bexsero, Trumenba, Nimenrix</li> <li>Invasive pneumococcal disease (“ENI” in Spanish)</li> <li>Influenza</li> <li>Hepatitis</li> <li>Rotavirus: Rotarix, Rotateq</li> <li>Measles (“Sarampión” in Spanish) and MMR (“Triple vírica” in Spanish)</li> <li>Sepsis</li> <li>Whooping cough (“Tosferina” in Spanish)</li> <li>Chickenpox (“Varicela” in Spanish): Varivax, Varilrix; and Shingles (“Zoster” in Spanish)</li> <li>Human papillomavirus infection (“VPH” in Spanish): Cervarix, Gardasil</li> </ul> <p>Tweets have been manually classified as having a negative or non-negative sentiment by 5 experts. Moreover, an automatic classification has been performed by 3 different tools: IBM Watson (now Watson Tone Analyzer, <a href="https://www.ibm.com/watson/services/tone-analyzer/">https://www.ibm.com/watson/services/tone-analyzer/</a>), Google Cloud Natural Language (<a href="https://cloud.google.com/natural-language">https://cloud.google.com/natural-language</a>), and Meaning Cloud (<a href="https://www.meaningcloud.com/">https://www.meaningcloud.com/</a>). IBM Watson and Google Cloud Natural Language returned a numerical sentiment score ranging from -1 to 1, while Meaning Cloud returned a categorical variable with the values ‘P+’, ‘P’, ‘NEU’, ‘N’ and ‘N+’, which were converted to 1, 2, 3, 4 and 5 respectively.</p> <p>With these variables (IBM Watson, Google Cloud Natural Language, and Meaning Cloud annotations and the experts’ classification as the target label), a machine learning metamodel was developed. Tweets were also annotated with the sentiment output given by this classifier. </p> <p>The provided data includes intrinsic tweets information, intrinsic information regarding the users that posted the tweets, the keywords mentioned in each tweet, and the annotations that the experts, the tools, and the model gave to each tweet.</p> <p><strong>Funding</strong>: This dataset was obtained with funding from MSD, Spain under MAVIS Study (VEAP ID: 7789).</p> <p><strong>Current studies using this dataset at the moment of the publication</strong>:</p> <ul> <li>Rodríguez-González et al., “Creating a metamodel based on machine learning to identify the sentiment of vaccine and disease-related messages in Twitter: the MAVIS study” in 2020 IEEE 33st International Symposium on Computer-Based Medical Systems (CBMS), Jul. 2020, p. 6. DOI: 10.1109/CBMS49503.2020.00053</li> <li>Rodríguez-González et al., "Identifying Polarity in Tweets from an Imbalanced Dataset about Diseases and Vaccines Using a Meta-Model Based on Machine Learning Techniques" in Applied Sciences, 2020, 10. DOI: 10.3390/app10249019</li> </ul>
A Multilingual Dataset of COVID-19 Vaccination Attitudes on Twitter
<p>This dataset consists of the IDs of 2,198,090 tweets collected from Western Europe, of which 17,934 are annotated with labels indicating the originators' affective vaccination stances, including Positive (PO), Negative (NG), Positive but dissatisfaction (PD), Neutral (NE) and Off-topic (OT).</p> <p>all_tweets.txt contains all the ids of the collected tweets, annotated_tweets.txt contains the ids of the annotated tweets and the categories they are annotated to.</p> <p> </p>
Twitter Dataset - Over 200,000 Tweets containing the word "Vaccine" for research porpuses
<p>This dataset contains 220,085 tweets containing the word vaccine between December 9th and December 18th 2021 at different times during each day, extracted using the Twitter API v2. Each tweet was extracted at least 3 days after its initial posting time in order to register 3 days of engagements, and it doesn't include retweets.</p> <p>Includes:</p> <ul> <li>Tweet ID</li> <li>Text</li> <li>Author ID</li> <li>Date</li> <li>Like count</li> <li>Retweet count</li> <li>Quote count</li> <li>Reply count</li> <li>User data (Followers, Following, Tweet count, Account creation date, Verified status)</li> </ul> <p>Usernames are hidden for privacy reasons</p>
Moral Values of Twitter COVID-19 Vaccine Data
<p>This data is part of our accepted paper "Learning to Adapt Domain Shifts of Moral Values via Instance Weighting" at the 33rd ACM Conference on Hypertext and Social Media (HT ’22). We annotate moral values of COVID-19 vaccine-related tweets. </p>
Attitude and CTM predictions for From Tribal Polarization to Socio-Economic Disparities: Exploring the Landscape of Vaccine Hesitancy on Twitter paper
<p>The presented data pertains to predictions of Attitudes and CTM, generated through the application of Machine Learning models. These models have been extensively elucidated in the research paper titled "From Tribal Polarization to Socio-Economic Disparities: Exploring the Landscape of Vaccine Hesitancy on Twitter". The aforementioned data is available to the public.</p>
English Vaccine-related Twitter data (Tweet IDs) 2018-07-01 to 2020-09-30
<p>Twitter data collected through the Crowdbreaks platform [1]. Between July 1st, 2017 and October 1st, 2020 a total of 57.5M tweets (including 39.7M retweets) in English language by 9.9M unique users were collected using the public filter stream endpoint of the Twitter API. The tweets matched one or more of the keywords "vaccine", "vaccination", "vaxxer", "vaxxed", "vaccinated", "vaccinating", "vacine", "overvaccinate", "undervaccinate", "unvaccinated". The data can be considered complete with respect to these keywords.</p> <p> </p> <p>The files contain a single column which corresponds to the Tweet ID (one file per day). Using the IDs the full tweet objects can be restored.</p>
Twitter vaccine misinformation data
<p>Anti-vaccine content is rapidly propagated via social media, fostering vaccine hesitancy, while pro-vaccine content has not replicated the opponent's successes. Despite this disparity in the dissemination of anti- and pro-vaccine posts, linguistic features that facilitate or inhibit the propagation of vaccine-related content remain less known. Moreover, most prior machine-learning algorithms classified social-media posts into binary categories (e.g., misinformation or not) and have rarely tackled a higher-order classification task based on divergent perspectives about vaccines (e.g., anti-vaccine, pro-vaccine, and neutral). Our objectives are (1) to identify sets of linguistic features that facilitate and inhibit the propagation of vaccine-related content and (2) to compare whether anti-vaccine, pro-vaccine, and neutral tweets contain either set more frequently than the others. To achieve these goals, we collected a large set of social media posts (over 120 million tweets) between Nov. 15 and Dec. 15, 2021, coinciding with the Omicron variant surge. A two-stage framework was developed using a fine-tuned BERT classifier, demonstrating over 99 and 80 percent accuracy for binary and ternary classification. Finally, the Linguistic Inquiry Word Count text analysis tool was used to count linguistic features in each classified tweet. Our regression results show that anti-vaccine tweets are propagated (i.e., retweeted), while pro-vaccine tweets garner passive endorsements (i.e., favorited). Our results also yielded the two sets of linguistic features as facilitators and inhibitors of the propagation of vaccine-related tweets. Finally, our regression results show that anti-vaccine tweets tend to use the facilitators, while pro-vaccine counterparts employ the inhibitors. These findings and algorithms from this study will aid public health officials' efforts to counteract vaccine misinformation, thereby facilitating the delivery of preventive measures during pandemics and epidemics.</p>
Twitter vaccine misinformation data
Open the record for dataset details and reuse information.
Hashtag HPV: HPV Vaccine Twitter Education Program
ClinicalTrials.gov study NCT05204030. IPD Sharing: NO. Countries: 1. Publications: 3.
Promoting HPV Vaccine Through Twitter
ClinicalTrials.gov study NCT04023955. IPD Sharing: NO. Countries: 0. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.