Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

11 results for “Twitter, vaccination”

Learn how ShareScore rates datasets ↗
zenodo48/100

VaxxHesitancy: A Dataset for Studying Hesitancy Towards COVID-19 Vaccination on Twitter

<p>We create a publicly available dataset of over 3,100 COVID-19 vaccine-related tweets labeled as one of four stance categories: <em>pro-vaxx, anti-vaxx</em>, <em>vaxx-hesitant</em>,<em> or irrelevant</em>.</p> <p><strong>***</strong></p> <p><strong>Please use the V2 version.</strong></p> <p><strong>***</strong></p> <p>We split our dataset into two separate files:</p> <p>(1) VaccineHesitancy_train_v2.csv (Single + Double annotated)</p> <p>(2) VaccineHesitancy_test.csv (Double annotated)</p> <p>We present the details of this dataset here:</p> <p>VaxxHesitancy: A Dataset for Studying Hesitancy Towards COVID-19 Vaccination on Twitter (ICWSM 2023)</p> <p><strong>Our Pre-trained model</strong> (GateNLP/covid-vaccine-twitter-bert) : https://huggingface.co/GateNLP/covid-vaccine-twitter-bert</p> <p><strong>Paper</strong>: https://ojs.aaai.org/index.php/ICWSM/article/view/22213/21992</p> <p>&nbsp;</p> <pre>@inproceedings{mu2023vaxxhesitancy, title={VaxxHesitancy: A Dataset for Studying Hesitancy Towards COVID-19 Vaccination on Twitter}, author={Mu, Yida and Jin, Mali and Grimshaw, Charlie and Scarton, Carolina and Bontcheva, Kalina and Song, Xingyi}, booktitle={Proceedings of the International AAAI Conference on Web and Social Media}, volume={17}, pages={1052--1062}, year={2023} } </pre> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

MAVIS Twitter dataset: A collection of tweets and sentiment analysis in Spanish about vaccines and diseases during the period 2015-2018

<p>MAVIS dataset comprises a full knowledge base regarding Twitter messages published in Spanish during the period 2015-2018, in the context of sentiment analysis of specific vaccines and their related diseases. Such diseases and vaccines are summarized as follows:</p> <ul> <li>Invasive meningococcal disease (&ldquo;EMI&rdquo; in Spanish): Bexsero, Trumenba, Nimenrix</li> <li>Invasive pneumococcal disease (&ldquo;ENI&rdquo; in Spanish)</li> <li>Influenza</li> <li>Hepatitis</li> <li>Rotavirus: Rotarix, Rotateq</li> <li>Measles (&ldquo;Sarampi&oacute;n&rdquo; in Spanish) and MMR (&ldquo;Triple v&iacute;rica&rdquo; in Spanish)</li> <li>Sepsis</li> <li>Whooping cough (&ldquo;Tosferina&rdquo; in Spanish)</li> <li>Chickenpox (&ldquo;Varicela&rdquo; in Spanish): Varivax, Varilrix; and Shingles (&ldquo;Zoster&rdquo; in Spanish)</li> <li>Human papillomavirus infection (&ldquo;VPH&rdquo; in Spanish): Cervarix, Gardasil</li> </ul> <p>Tweets have been manually classified as having a negative or non-negative sentiment by 5 experts. Moreover, an automatic classification has been performed by 3 different tools: IBM Watson (now Watson Tone Analyzer, <a href="https://www.ibm.com/watson/services/tone-analyzer/">https://www.ibm.com/watson/services/tone-analyzer/</a>), Google Cloud Natural Language (<a href="https://cloud.google.com/natural-language">https://cloud.google.com/natural-language</a>), and Meaning Cloud (<a href="https://www.meaningcloud.com/">https://www.meaningcloud.com/</a>). IBM Watson and Google Cloud Natural Language returned a numerical sentiment score ranging from -1 to 1, while Meaning Cloud returned a categorical variable with the values &lsquo;P+&rsquo;, &lsquo;P&rsquo;, &lsquo;NEU&rsquo;, &lsquo;N&rsquo; and &lsquo;N+&rsquo;, which were converted to 1, 2, 3, 4 and 5 respectively.</p> <p>With these variables (IBM Watson, Google Cloud Natural Language, and Meaning Cloud annotations and the experts&rsquo; classification as the target label), a machine learning metamodel was developed. Tweets were also annotated with the sentiment output given by this classifier. &nbsp;&nbsp;</p> <p>The provided data includes intrinsic tweets information, intrinsic information regarding the users that posted the tweets, the keywords mentioned in each tweet, and the annotations that the experts, the tools, and the model gave to each tweet.</p> <p><strong>Funding</strong>: This dataset was obtained with funding from&nbsp;MSD, Spain under MAVIS Study (VEAP ID: 7789).</p> <p><strong>Current studies using this dataset at the moment of the publication</strong>:</p> <ul> <li>Rodr&iacute;guez-Gonz&aacute;lez et al., &ldquo;Creating a metamodel based on machine learning to identify the sentiment of vaccine and disease-related messages in Twitter: the MAVIS study&rdquo; in 2020 IEEE 33st International Symposium on Computer-Based Medical Systems (CBMS), Jul. 2020, p. 6. DOI: 10.1109/CBMS49503.2020.00053</li> <li>Rodr&iacute;guez-Gonz&aacute;lez et al., &quot;Identifying Polarity in Tweets from an Imbalanced Dataset about Diseases and Vaccines Using a Meta-Model Based on Machine Learning Techniques&quot; in Applied Sciences, 2020, 10. DOI: 10.3390/app10249019</li> </ul>

opencc-by-4.0Dec 2020View details →
zenodo44/100

A Multilingual Dataset of COVID-19 Vaccination Attitudes on Twitter

<p>This dataset consists of the IDs of 2,198,090 tweets collected from Western Europe,&nbsp;of which 17,934 are annotated with&nbsp;labels indicating the originators&#39; affective vaccination stances,&nbsp;including Positive (PO), Negative (NG), Positive but dissatisfaction (PD), Neutral (NE) and Off-topic (OT).</p> <p>all_tweets.txt contains all the ids of the collected tweets, annotated_tweets.txt&nbsp;contains the ids of the annotated tweets and the categories they are annotated to.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Twitter Dataset - Over 200,000 Tweets containing the word "Vaccine" for research porpuses

<p>This dataset contains 220,085 tweets containing the word vaccine between December 9th and December 18th 2021 at different times during each day, extracted using the Twitter API v2. Each tweet was extracted at least 3 days after its initial posting time in order to register 3 days of engagements, and it doesn&#39;t include retweets.</p> <p>Includes:</p> <ul> <li>Tweet ID</li> <li>Text</li> <li>Author ID</li> <li>Date</li> <li>Like count</li> <li>Retweet count</li> <li>Quote count</li> <li>Reply count</li> <li>User data (Followers, Following, Tweet count, Account creation date, Verified status)</li> </ul> <p>Usernames are hidden for privacy reasons</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Moral Values of Twitter COVID-19 Vaccine Data

<p>This data is part of our accepted paper &quot;Learning to Adapt Domain Shifts of Moral Values via Instance Weighting&quot; at&nbsp;the 33rd ACM Conference on Hypertext and Social Media (HT &rsquo;22). We annotate moral values of COVID-19 vaccine-related tweets.&nbsp;</p>

opencc-by-3.0-usApr 2022View details →
zenodo40/100

Attitude and CTM predictions for From Tribal Polarization to Socio-Economic Disparities: Exploring the Landscape of Vaccine Hesitancy on Twitter paper

<p>The presented data pertains to predictions of Attitudes and CTM, generated through the application of Machine Learning models. These models have been extensively elucidated in the research paper titled &quot;From Tribal Polarization to Socio-Economic Disparities: Exploring the Landscape of Vaccine Hesitancy on Twitter&quot;. The aforementioned data is available to the public.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

English Vaccine-related Twitter data (Tweet IDs) 2018-07-01 to 2020-09-30

<p>Twitter data collected through the Crowdbreaks platform [1]. Between July 1st, 2017 and October 1st, 2020 a total of 57.5M tweets (including 39.7M retweets) in English language by 9.9M unique users were collected using the public filter stream endpoint of the Twitter API. The tweets matched one or more of the keywords &quot;vaccine&quot;, &quot;vaccination&quot;, &quot;vaxxer&quot;, &quot;vaxxed&quot;, &quot;vaccinated&quot;, &quot;vaccinating&quot;, &quot;vacine&quot;, &quot;overvaccinate&quot;, &quot;undervaccinate&quot;, &quot;unvaccinated&quot;. The data can be considered complete with respect to these keywords.</p> <p>&nbsp;</p> <p>The files contain a single column which corresponds to the Tweet ID (one file per day). Using the IDs the full tweet objects can be restored.</p>

opencc-by-4.0Nov 2020View details →
dryad36/100

Twitter vaccine misinformation data

<p>Anti-vaccine content is rapidly propagated via social media, fostering vaccine hesitancy, while pro-vaccine content has not replicated the opponent's successes. Despite this disparity in the dissemination of anti- and pro-vaccine posts, linguistic features that facilitate or inhibit the propagation of vaccine-related content remain less known. Moreover, most prior machine-learning algorithms classified social-media posts into binary categories (e.g., misinformation or not) and have rarely tackled a higher-order classification task based on divergent perspectives about vaccines (e.g., anti-vaccine, pro-vaccine, and neutral). Our objectives are (1) to identify sets of linguistic features that facilitate and inhibit the propagation of vaccine-related content and (2) to compare whether anti-vaccine, pro-vaccine, and neutral tweets contain either set more frequently than the others. To achieve these goals, we collected a large set of social media posts (over 120 million tweets) between Nov. 15 and Dec. 15, 2021, coinciding with the Omicron variant surge. A two-stage framework was developed using a fine-tuned BERT classifier, demonstrating over 99 and 80 percent accuracy for binary and ternary classification. Finally, the Linguistic Inquiry Word Count text analysis tool was used to count linguistic features in each classified tweet. Our regression results show that anti-vaccine tweets are propagated (i.e., retweeted), while pro-vaccine tweets garner passive endorsements (i.e., favorited). Our results also yielded the two sets of linguistic features as facilitators and inhibitors of the propagation of vaccine-related tweets. Finally, our regression results show that anti-vaccine tweets tend to use the facilitators, while pro-vaccine counterparts employ the inhibitors. These findings and algorithms from this study will aid public health officials' efforts to counteract vaccine misinformation, thereby facilitating the delivery of preventive measures during pandemics and epidemics.</p>

opencc-zeroJun 2022View details →
dryad36/100

Twitter vaccine misinformation data

Open the record for dataset details and reuse information.

publicDec 2022View details →
ClinicalTrials.gov32/100

Hashtag HPV: HPV Vaccine Twitter Education Program

ClinicalTrials.gov study NCT05204030. IPD Sharing: NO. Countries: 1. Publications: 3.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov20/100

Promoting HPV Vaccine Through Twitter

ClinicalTrials.gov study NCT04023955. IPD Sharing: NO. Countries: 0. Publications: 0.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record