Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

307

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

307 results for “Twitter”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dataset for supervised bot detection on Twitter

<p>This is the dataset I used for my Bachelor Thesis &quot;Fake-tweet: Desarrollo de un modelo para la detecci&oacute;n de bots en Twitter&quot;, where I developed a supervised method for bot detection on Twitter.</p> <p>This&nbsp;dataset contains both genuine users&nbsp;(3474) and social spambots (4912)&nbsp;and has 69 features defined for each of these accounts. Such features can be divided&nbsp;in 3 categories:&nbsp;content,&nbsp;account information or account usage features.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Graphic for Twitter - Call for Open Science Success Stories

<p>This graphic was created to share on Twitter to encourage people to submit compelling stories about open science practices in research to the TOPS team. This graphic was tweeted out on Dr. Chelle Gentemann&#39;s Twitter account on August 23, 2022. <a href="https://twitter.com/ChelleGentemann/status/1562049765566709764?s=20&amp;t=kkxyz4dx1vA6obyveL5WNA">Link to tweet</a>.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Datasets containing the results from the analysis on SDGS and eHealth inside the Citizen Science Community on Twitter

<p>This datasets contain the results from our analyses of the Citizen Science Community on Twitter. These analyses have been done to better understand the discussion about SDGs, eLearning&nbsp;and eHealth.</p> <p><strong>T</strong>he purpose of sharing these datasets&nbsp;is to provide the basis to reproduce&nbsp;the results reported in the associated deliverable. These files are not raw data, since due to privacy concerns we can not share personal information from Twitter.</p> <p><strong>dominant_topics_anonym.xlsx</strong>: Excel datasheet. This dataset contians the distribution of the most discussed topics inside the SDGs discussion.</p> <p><strong>Edges_Hashtag_connected.csv</strong>:&nbsp;&nbsp;CSV file. This dataset contains the edges to build the network of connected hashtags.&nbsp;This edges can be used to build a network and explore the connections or to statiscally analyse the results.</p> <p><strong>hashtags.csv</strong>: CSV file. This dataset contains the results of the most used hashtags in the analysis about eLearning.&nbsp;<br> &nbsp;</p> <p><strong>hashtags_treemap_health.xlsx</strong>: Excel datasheet. This dataset contains the results of the most frequent hashtags in the eHealth analysis.</p> <p><strong>ldavis_prepared_ieee17.html</strong>: HTML file. This file contains the Intertopic distance map and most salient terms from the topic modelling analysis done in the SDGs conversation study.</p> <p><strong>Most_retweeted_accounts.xlsx</strong>: Excel datasheet. This dataset contains the top 20 users that receive more retweets in the conversation around eHealth. The column called&nbsp;Indegree refers to the topological value calculated from the network of retweets. This indegree is equivalent to the number of retweets received. On the other hand, Outdegree is the opposite, so number of retweets given to others.</p> <p><strong>Most_retweeting_account.xlsx</strong>: Excel datasheet. This dataset presents the opposite part of the previous one, the accounts that retweet the most from the eHealth analysis. The columns contain the same indicators: Indegree and Outdegree.</p> <p><strong>sdgs_count_publish.csv</strong>: CSV file. This dataset contains the number of tweets assigned to the different SDGs from the analysis done on the conversation about these Goals.</p> <p><strong>sdgs_tweets_sdgsaccess.xlsx</strong>: Excel datasheet. Same file as the previous one in other format to ease the handling in Excel.</p> <p><strong>top_hash_health.xlsx</strong>: Excel datasheet. The most used hashtags inside the conversation about eHealth.</p> <p><strong>topics_tweets_sdgsaccess.xlsx</strong>: Excel datasheet. Tweets by topic extracted using Machine Learning in the SDGs analysis.</p> <p>&nbsp;</p> <p>This repository will receive updates in the future in order to present all the data available and publishable from the different analysis that were described.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Tagged Twitter timelines for users reporting SARS-CoV-2 infections and related data

<p>Twitter data was collected through the Twitter API v2.0, specifically through the timeline endpoint. Details about the inference of SARS-CoV-2 self-reports and the tagging of the full timeline of each user can be found in the <a href="https://github.com/digitalepidemiologylab/content_changes_paper">GitHub repository</a>. The larger dataset (&quot;preprocessed_data.csv&quot;) consists of a total of 8,534,171 tweets posted by 30,856&nbsp;users from January 1, 2020 to October to September 30, 2021.</p> <p>The raw data from Twitter, including tweet and user IDs, has been removed or anonymized in order to comply with the EPFL guidelines for data sharing.</p> <p>In particular, the date of the tweets was removed, the text&nbsp;of the tweets, URLs and URL domains&nbsp;have been substituted with the &quot;text&quot;, &quot;&lt;URL&gt;&quot; and &quot;&lt;URL_DOMAIN&gt;&quot; token respectively.</p> <p>User IDs have been substitued with new IDs in the [0, number of users] range (e.g. U0, U1, ...) .</p> <p>Tweet IDs have been substitued with new IDs in the [0, number of tweets] range (e.g. T0, T1, ...) .</p> <p>In addition to self-explanatory columns about topics, emotions, URL classification and symptoms we tagged, we also share the columns:</p> <ul> <li>pdate: date of the SARS-CoV-2 infection self-report for that user (adjusted with SUTime)</li> <li>effective_date: date of the tweet adjusted with SUTime, when the SUTime columns is available.</li> <li>rel_effective_day(week, month): days (weeks, months) computed with respect to the positivity date (i.e. &quot;pdate&quot; column). Negative numbers refer to tweets posted before the user reported a COVID-19 infection on Twitter.</li> </ul>

opencc-by-4.0Jan 2023View details →
zenodo44/100

TweetDIS: A Large Twitter Dataset for Natural Disasters Built using Weak Supervision

<p>This repository contains the silver standard dataset and code for the paper &quot;TweetDIS: A Large Twitter Dataset for Natural Disasters Built using Weak Supervision&quot;.</p> <p>The file &quot;heuristic_uniq_terms_nd.txt&quot; contains the list of terms used as the heuristic and the file &quot;natural_disasters_ssd_tweetids.tsv&quot; contains the tweet ids in&nbsp;the silver standard dataset.&nbsp;</p> <p>To hydrate the tweets, you can use tools like twarc or Social Media Mining toolkit - https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7362951/</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Twitter Dataset for Pakistani Political Discourse

<p>This is one of the largest dataset of Pakistani Twitter political discourse and consists of more than 49 million tweets, collected during&nbsp; April 2022.</p> <p>Please use the following citation if you use this this dataset: MLA format: Haq, Ehsan-Ul, et al. &quot;A Twitter Dataset for Pakistani Political Discourse.&quot; arXiv preprint arXiv:<a href="https://doi.org/10.48550/arXiv.2301.06316">2301.06316</a> (2023), doi:&nbsp;<a href="https://doi.org/10.48550/arXiv.2301.06316">https://doi.org/10.48550/arXiv.2301.06316</a></p> <p>Bibtex: @misc{2301.06316, Author = {Ehsan-Ul Haq and Haris Bin Zia and Reza Hadi Mogavi and Gareth Tyson and Yang K. Lu and Tristan Braud and Pan Hui}, Title = {A Twitter Dataset for Pakistani Political Discourse}, Year = {2023}, Eprint = {arXiv:2301.06316}, }</p> <p>Relevant paper highlight the dataset collection, and any possible changes is accessible at: https://arxiv.org/abs/2301.06316</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Code-mixed Indonesian-Javanese-English Twitter Dataset

<p>This is a Twitter dataset for code-mixed language identification. The dataset contains mixed Indonesian, Javanese, and English words.&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

A dataset of gamers on Twitter

<p>This gaming-related dataset consists of 8932 users (labeled as&nbsp;<em>gamers</em>) engaging in game-related conversations. We have collected (June 2018) their timeline (the most recent 3200 tweets) using the Twitter Search API.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

A dataset of journalists on Twitter

<p>This dataset comprises the Twitter timelines of journalists belonging to 17 different countries from 8 different continental regions, downloaded&nbsp;in May 2018. We used the Twitter REST API, which allows us to download, for a given user, the last 3200 tweets since the time of the download.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Data and code for the paper "Attention, sentiments and emotions towards emerging climate technologies on Twitter"

<p>This is the code and data for the paper "Attention, sentiments and emotions towards emerging climate technologies on Twitter" by Müller-Hansen et al. (Global Environmental Change, 2023).</p><p>This archive contains the following materials:</p><ul><li>Data set of tweets</li><li>Table of subqueries for searching Twitter</li><li>Code for figure generation</li></ul><p>Please see the Readme for further details.</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Covid-19 Twitter Database: Turkish Sample

<p>This dataset is a comprehensive dataset that includes&nbsp;Turkish tweet ids one month before and after of pandemic outbreak. You can see the tweets classified across COVID, economy, politics, religion, disinformation, international relations&nbsp;themes.&nbsp;To see research design and analyze the code for the text analysis please visit this link:&nbsp;<a href="https://github.com/burakozturan/tria-covid19">https://github.com/burakozturan/css_covid19</a></p>

opencc-by-4.0Jun 2020View details →
zenodo40/100

TweetBLM: A Hate Speech Dataset and Analysis of BlackLivesMatter-related Microblogs on Twitter

<p>Collection of BLM related tweets and their corresponding labels of hate speech.</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Partidos_Politicos_Twitter

<p>El data set est&aacute; compuesto por nueve variables. Semana que hace referencia a la fecha semanal de los datos de las otras variables. Seguidores del partido (hay una por cada partido), que son el n&uacute;mero de nuevos seguidores que ha conseguido el partido esa semana. Seguidos del partido (hay una por cada partido), que indica el n&uacute;mero de cuentas nuevas que sigue el partido en Twitter, este puede ser negativo en caso de dejar de seguir cuentas.&nbsp; El data set contiene los datos del PSOE, PP, Podemos y Vox.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Datasets of Twitter mentions and publications in Information Science & Library Science and Microbiology

<p>Datasets used in the study &#39;Identifying and characterizing social media communities: a socio-semantic network approach to altmetrics&#39;.</p> <p><strong>Microbiology publications (mic_publiccations.tsv).</strong> Dataset of 101,206 Microbiology publications with their author keywords.</p> <p><strong>Microbiology mentions (mic_mentions.tsv).</strong> Dataset of 328,110 Twitter mentions to Microbiology publications.</p> <p><strong>Information Science &amp; Library Science publications (lis_publications.tsv).</strong> Dataset of 8452 Information Science &amp; Library Science publications with their author keywords.</p> <p><strong>Information Science &amp; Library Science mentions (lis_mentions.tsv).</strong> Dataset of 35,411 Twitter mentions to Information Science &amp; Library Science publications.</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Dataset for English Health-Related Advice Directed to the General Public on Twitter During the Early Spread of COVID-19 [Dataset]

<p>Health-related advice directed to the public on twitterprovides insight into the use of social media duringa pandemic. This paper describes our data collection, sampling, and analysis of 44 million tweets in English in March 2020. We make reference to a parallel dataset and analysis of tweets in Arabic during thesame period. The contribution of this paper is a description of our dataset, our coding process to indicate tweets with health related advice, and our analysis and comparisons of the characteristics of the tweets with and without health-related advice. These contributions providethe basis for future research on semi-automated classifiers for health-related advice and efforts to reduce thespread of harmful health advice.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

AraHealth: A Dataset for Arabic Health-Related Advice Directed to the General Public on Twitter During the Early Spread of COVID-19 [Dataset]

<p>Health-related advice directed to the general public on Twitter provides insight into the use of social media during health emergencies. This paper describes our data collection, sampling, and analysis of 24 million tweets in Arabic in March and early April 2020. We make reference to a parallel dataset and analysis of tweets in English during the same period. The contribution of this paper is a description of our dataset, our coding process to indiciate tweets with health related advice, and our analysis and comparisons of the characteristics of the tweets with and without health-related advice. These contributions provide the basis for future research on semi-automated classifiers for health-related advice and efforts to reduce the spread of harmful health advice.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

A Geo-Tagged COVID-19 Twitter Dataset for 10 North American Metropolitan Areas

<p>The dataset comprises of 10 JSON files, each containing geographic metadata and a sentiment score collected from tweets between March 20, 2020 and December 1, 2020 pertaining to the COVID-19 global pandemic for ten of the most populous cities in the United States and Canada.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Twitter topics list

<p>List of <a href="https://blog.twitter.com/en_us/topics/product/2019/introducing-topics.html">Topics</a> on Twitter as of 1/19/2021.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Following/Followers and Tags on 0.1 million Twitter Users

<p><strong>Abstract</strong> (our paper)</p> <p>Why does Smith follow Johnson on Twitter? In most cases, the reason why users follow other users is unavailable. In this work, we answer this question by proposing TagF, which analyzes the who-follows-whom network (matrix) and the who-tags-whom network (tensor) simultaneously. Concretely, our method decomposes a coupled tensor constructed from these matrix and tensor. The experimental results on million-scale Twitter networks show that TagF uncovers different, but explainable reasons why users follow other users.</p> <p><strong>Data</strong></p> <p>coupled_tensor:<br> The first column is the source user id (from user id), the second column is the destination user id (to user id), and the third column is the tag id.</p> <p>users.id:<br> The first column is the user id for coupled_tensor, and the second column is the user id on Twitter.</p> <p>tags.id:<br> The first column is the tag id for coupled_tensor, and the second column is the tag (<em>i.e.</em> slug or list name) on Twitter. On the tags, ###follow### and ###friend### are special tags expressing follower and following.</p> <p><strong>Publication</strong></p> <p>This dataset was created for our study. If you make use of this dataset, please cite:<br> Yuto Yamaguchi, Mitsuo Yoshida, Christos Faloutsos, Hiroyuki Kitagawa. Why Do You Follow Him? Multilinear Analysis on Twitter. <em>Proceedings of the 24th International Conference on World Wide Web (WWW '15 Companion)</em>. pp.137-138, 2015.<br> http://doi.org/10.1145/2740908.2742715</p> <p><strong>Code</strong></p> <p>Our code outputting experiment results made available at:<br> https://github.com/yamaguchiyuto/tagf</p> <p><strong>Note</strong></p> <p>If you would like to use larger dataset, the dataset on 1 million seed users made available at:<br> http://dx.doi.org/10.5281/zenodo.16267<br> (The dataset on 0.1 million seed users is not subset of the dataset on 1 million seed users.)</p>

opencc-zeroJan 2015View details →
zenodo40/100

Can accurate demographic information about people who use prescription medications non-medically be derived from Twitter?

<p>This archive contains over 3 billion Tweet IDs associated with the paper:<br>"Can accurate demographic information about people who use prescription medications non-medically be derived from Twitter?"</p> <p>The data provides:<br>X (Twitter) IDs of posts included in the study. The IDs can be used to retrieve the original posts via the Twitter API. Posts removed by the original subscribers or whose visibilities are no longer public cannot be retrieved by the poster.&nbsp;</p> <p>Full citation:<br>Yang YC, Al-Garadi MA, Love JS, Cooper HLF, Perrone J, Sarker A. Can accurate demographic information about people who use prescription medications nonmedically be derived from Twitter? Proc Natl Acad Sci U S A. 2023 Feb 21;120(8):e2207391120. doi: 10.1073/pnas.2207391120. Epub 2023 Feb 14. PMID: 36787355; PMCID: PMC9974473.</p> <p>Python scripts related to the analysis are available as supplementary material with the paper.&nbsp;</p> <p>Contact:&nbsp;<br>Abeed Sarker<br>abeed@dbmi.emory.edu</p> <p>Funding:<br>National Institute on Drug Abuse (R01DA057599).</p>

opencc-by-4.0Dec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record