Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

307

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

307 results for “Twitter”

Learn how ShareScore rates datasets ↗
zenodo40/100

'Twitter Is Dead, Long Live X!' A Decade of Microblog Research and Implications for Knowledge and Research

<p>Talk at the <a href="https://www.digital-philosophy.org/" target="_blank" rel="noopener">Philosophy [in:of:for:and] Digital Knowledge Infrastructures</a> online workshop 2023 (28/09/2023).</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Open dataset of scholars on Twitter (X)

<p>This is a version 2 dataset of paired OpenAlex author IDs (<a href="https://docs.openalex.org/about-the-data/author">https://docs.openalex.org/about-the-data/author</a>) and Twitter (now X) user IDs<br><br></p> <p><strong>Major update in this version</strong></p> <p>Following the significant update to OpenAlex's author identification system, the scholars on Twitter dataset, which previously linked Twitter IDs to OpenAlex author IDs, immediately became outdated. This called for a new approach to re-establish these links, as the absence of new Twitter data made it impossible to replicate the original method of matching Twitter profiles with scholarly authors. To navigate this challenge, a bridge was constructed between the June 2022 snapshot of the OpenAlex database&mdash;used in the original matching process&mdash;and the most recent snapshot from February 2024. This bridge utilized OpenAlex works IDs and DOIs to match authors in both datasets by their shared publications and identical primary names. When a connection was established between two authors with the same name, the new OpenAlex author ID was assigned to the corresponding Twitter ID. When direct matches based on primary names were not found, an attempt was made to establish connections by matching the names from June 2022 with any corresponding alternative names found in the 2024 dataset. This method ensured continuity of identity through the system update, adapting the strategy to link profiles across the temporal divide created by the database's overhaul.</p> <p>Our efficient method for re-establishing links between author IDs and Twitter profiles has been notably successful, managing to rematch 432,417 (88%) OpenAlex author IDs. This effort successfully restored connections for 388,968 unique Twitter users, which represents 92% of the original dataset. Of these, 375,316 were matched using their primary names, and 57,101 through alternative names. The simplicity and quick execution of this approach led to exceptionally favourable results, with a minimal loss of only 8% of the original Twitter-linked scholarly accounts.</p> <p>The dataset includes <strong>&nbsp;432,417 unique author_ids</strong> and&nbsp;<strong>388,968 unique tweeter_ids</strong> forming <strong>462,427 unique author-tweeter pairs</strong>.</p> <p>&nbsp;</p> <p><strong>File descriptions</strong></p> <ul> <li><strong>authors_tweeters_2024_02.csv</strong> is the actual dataset of author IDs paired with tweeter IDs. The "alternative" column indicates if the match was made with the primary name (0) or an alternate name (1).</li> <li><strong>mapping_tweeters_2022_2024.csv</strong> contains the relationship made between the 2022 author IDs and the 2024 author IDs, including the names.</li> </ul> <p>&nbsp;</p> <p><strong>How to cite</strong></p> <p>When using the dataset, please cite the following article providing details about the matching process:</p> <div> <div>Mongeon, P., Bowman, T. D., &amp; Costas, R. (2023). An open data set of scholars on Twitter. <em>Quantitative Science Studies</em>, 1&ndash;11. <div><a href="https://doi.org/10.1162/qss_a_00250">https://doi.org/10.1162/qss_a_00250</a></div> </div> </div>

opencc-zeroMar 2024View details →
zenodo40/100

List of Twitter bots that mention scientific publications

<p>This collection of datasets comes from the paper titled <em>"The Botization of Science? Large-scale study of the presence and impact of Twitter bots in science dissemination"</em> and includes two files:</p> <ol> <li><strong>Twitter_bots_list.txt</strong> - Contains a list of 11,073 Twitter bots that mention scientific publications, as identified in the study.</li> <li> <p><strong>Twitter_botscores.tsv</strong> - Includes the botscores of 4,872,369 Twitter accounts analyzed in the study, providing a measure of the likelihood that an account is a bot.</p> </li> </ol>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Recull d'opinions sobre els efectes del Brexit a Twitter.

<p>En aquest dataset trobarem un recull d&rsquo;opinions per part d&rsquo;usuaris de la xarxa social Twitter sobre el Brexit. D&rsquo;aquesta forma trobar quina relaci&oacute; observa la poblaci&oacute; entre el Brexit i els esdeveniments del present. En diferents camps hi trobarem tant les opinions com la id i la data del tuit. Els tuits s&oacute;n comentaris trobats com a resultat de la cerca de la paraula &ldquo;brexit&rdquo; en el cercador i filtrant els 1000 primers resultats. Les 1000 entrades que tindrem seran les m&eacute;s recents a data d&#39;extracci&oacute; per&ograve; com la data de la publicaci&oacute; del tweet sortir&agrave; al dataset, sempre es podr&agrave; contextualitzar.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

TNCD (Twitter News Cascade Dataset)

<p>The TNCD (Twitter News Cascade Dataset) 1,200 news items posted on Twitter. It contains post URLs retrieved from sportsmen, politicians and news channels accounts, most from September 2020. For more information<br> please see&nbsp;<a href="https://www.sciencedirect.com/science/article/pii/S1568494621003367">https://www.sciencedirect.com/science/article/pii/S1568494621003367</a></p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Mensajes de twitter en la campaña electoral de las elecciones autonomicas de Madrid.

<p>Mensajes de twitter en la campa&ntilde;a electoral de las elecciones autonomicas de Madrid.</p> <p>Se presentan los tuits en campa&ntilde;a electoral (1 mes antes de la jornada electoral)&nbsp; de los principales representantes politicos.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Fig. 4 in Ameronothrus twitter sp. nov. (Acari, Oribatida) a New Coastal Species of Oribatid Mite from Japan

Fig. 4. Ameronothrus twitter Pfingstl and Shimano, sp. nov., collecting place and macro photography images (A and B taken by SS; C, D and E original photographs shared on Twitter by TN). A, Collecting place on the wharf of the fishing port (Choshi-Gaikou), Chiba, Japan, arrow pointing to exact location of mites; B, location of a colony in a crack between a car chock and the wharf, arrow pointing to the exact location of the colony; C, a colony of A. twitter Pfingstl and Shimano, sp. nov. in the colder day (11:23 am), 5 May 2019 (https://twitter.com/ yatsume_project/status/1124999187626483717?s=20); D (https://twitter.com/yatsume_project/status/1123890470285918209?s=20) and E, moving individuals of A. twitter Pfingstl and Shimano, sp. nov. in the warmer day (11:57 am), 2 May 2019.

opencc-by-4.0Mar 2021View details →
zenodo40/100

Fig. 3 in Ameronothrus twitter sp. nov. (Acari, Oribatida) a New Coastal Species of Oribatid Mite from Japan

Fig. 3. Ameronothrus twitter Pfingstl and Shimano, sp. nov., adult legs. A, Right leg I dorsal view; B, left leg II dorso-lateral view of antiaxial aspect; C, left leg III antiaxial view; D, left leg IV antiaxial view. Scale bar valid for all legs.

opencc-by-4.0Mar 2021View details →
zenodo40/100

Fig. 2 in Ameronothrus twitter sp. nov. (Acari, Oribatida) a New Coastal Species of Oribatid Mite from Japan

Fig. 2. Ameronothrus twitter Pfingstl and Shimano, sp. nov., stereomicroscopic photographs of adult female. A, Dorsal view; B, ventral view; C, lateral view.

opencc-by-4.0Mar 2021View details →
zenodo40/100

Fig. 1 in Ameronothrus twitter sp. nov. (Acari, Oribatida) a New Coastal Species of Oribatid Mite from Japan

Fig. 1. Ameronothrus twitter Pfingstl and Shimano, sp. nov., adult female. A, Dorsal view, legs omitted; B, ventral view, mouthparts and legs distal segments omitted. Scale bar valid for both depictions.

opencc-by-4.0Mar 2021View details →
zenodo40/100

Portugues Twitter Covid-19 - mar/2020 & mar/2021

<p>Este &eacute; um dataset de tu&iacute;tes &uacute;nicos em portugu&ecirc;s relacionados &agrave;&nbsp;COVID-19.</p> <p>Os tu&iacute;tes compartilhados aqui possuem apenas o ID, devido aos termos e condi&ccedil;&otilde;es do Twitter para redistribuir dados do Twitter <strong>APENAS </strong>para prop&oacute;sito de pesquisa. Eles precisam ser hidratados para ser usados.</p> <p>Os conjuntos de dados cont&eacute;m tu&iacute;tes dos meses de mar&ccedil;o de 2020 e 2021. Ap&oacute;s a hidrata&ccedil;&atilde;o, ser&atilde;o 936.866 tu&iacute;tes de mar/2020 e 599.638 tu&iacute;tes de mar/ 2021.</p> <p>Este conjunto &eacute; um subproduto do trabalho de Banda&nbsp;<em>et al.</em>&nbsp;(2021). Link:&nbsp;https://zenodo.org/record/4603998#.YbkUIb3MJPa</p> <p>--------------</p> <p>This is a dataset with portuguese&nbsp;unique tweets related to COVID-19.</p> <p>Tweets shared here only have the ID, due to Twitter&#39;s terms and conditions to redistribute Twitter data for research purposes <strong>ONLY</strong>. They need to be hydrated to be used.</p> <p>The datasets contain tweets for the months of March 2020 and 2021. After hydration, there will be 936,866 tweets from Mar/2020 and 599,638 tweets from Mar/2021.</p> <p>This dataset is a by-product of the work by Banda et al. (2021).&nbsp;Link:&nbsp;https://zenodo.org/record/4603998#.YbkUIb3MJPa</p>

openother-pdDec 2021View details →
zenodo40/100

Twitter Dataset - Over 200,000 Tweets containing the word "Vaccine" for research porpuses

<p>This dataset contains 220,085 tweets containing the word vaccine between December 9th and December 18th 2021 at different times during each day, extracted using the Twitter API v2. Each tweet was extracted at least 3 days after its initial posting time in order to register 3 days of engagements, and it doesn&#39;t include retweets.</p> <p>Includes:</p> <ul> <li>Tweet ID</li> <li>Text</li> <li>Author ID</li> <li>Date</li> <li>Like count</li> <li>Retweet count</li> <li>Quote count</li> <li>Reply count</li> <li>User data (Followers, Following, Tweet count, Account creation date, Verified status)</li> </ul> <p>Usernames are hidden for privacy reasons</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

A Twitter dataset (with labels) for Life-Event detection

<p>The file contains an anonymized version of the dataset collected in the framework of the &ldquo;Tsundoku&rdquo; project financed by the Autonomous Province of Trento according to the province law 13 of December 1999, n. 6 (and subsequent modifications), art. 5. - &ldquo;Aids for promoting research and development&rdquo;, financing approved with APIAE manager&rsquo;s provision n. 691. The purpose of this project is training a Deep Learning (DL) model capable of detecting the occurrence of a so-called Life Event - a wedding and/or the birth of a child - in a person&rsquo;s life on the basis of the contents she shared on social media (Twitter, in this case).&nbsp;</p> <p>More precisely, the dataset consists of the most recent tweets - up to approximately 3200 for each account - of 8 Italian and 27 English-speaking users (all of them randomly picked), totalling 74722 tweets, 20302 written in Italian and 54420 in English.</p> <p>For each user (labelled as &lsquo;user_<em>x</em>&rsquo; with <em>x</em> an integer number between 1 and 35), the file includes the language her tweets are written in as well as the list of said tweets. For every one of the latter, the information made available includes the text of the tweet (appropriately modified as explained below), its length, the number of hashtags and of user mentions it featured, the number of retweets and of likes it received, two Boolean flags (&ldquo;True&rdquo; or &ldquo;False&rdquo;) assessing whether it was a quote or a reply and a label (&lsquo;birth&rsquo;, &lsquo;wedding&rsquo; and &lsquo;not Life-Event-related&rsquo;, depending on whether the tweet refers to a birth/wedding experienced by the user or not). With respect to the label, it is worth stressing that a tweet was labelled as &lsquo;birth&rsquo;/&lsquo;wedding&rsquo; only if the event &ldquo;actively&rdquo; involved the user (thus, a tweet reading &ldquo;Today I get married&rdquo; is labelled as &lsquo;wedding&rsquo;, while the label &lsquo;not Life-Event-related&rsquo; is associated with &lsquo;Today my sister gets married&rsquo;) and if the tweet is <strong>by itself</strong> unambiguous (in other words, a tweet reading &lsquo;My wife is pregnant&rsquo; is labelled as &lsquo;birth&rsquo;, while &lsquo;Josephine is pregnant&rsquo; does not, since the tweet alone does not allow to determine who Josephine exactly is - this could perhaps be inferred from other tweets but this kind of contextualization is very hard to be carried out and was out of the scope of this project.).</p> <p>In order to comply with GDPR, each one of the texts included in the file was obtained from the original text after taking the following anonymization steps:</p> <p>&nbsp;</p> <ul> <li> <p>every web link got replaced by the &lsquo;WEBLINK&rsquo; string;</p> </li> <li> <p>every mention to another Twitter user (for instance, &ldquo;@joedoe&rdquo;) got replaced by the &ldquo;OTHERUSER&rdquo; string;</p> </li> <li> <p>every hashtag got replaced by the &ldquo;HASHTAG&rdquo; string;</p> </li> <li> <p>every name and surname detected by the spaCy library (see <a href="https://spacy.io/">https://spacy.io</a> for more info) got replaced by the &ldquo;NAME/SURNAME&rdquo; string;</p> </li> <li> <p>10% of the available words got randomly picked and erased.&nbsp;&nbsp;</p> </li> </ul> <p>&nbsp;&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Twitter Timelines of British MPs (Tweet IDs)

<p>TSV file containing Twitter tweet_ids from the Timelines of 584&nbsp;members of British Parliament (collected between 4th and 6th of March 2022). The users were identified&nbsp;from the&nbsp;link below:</p> <p>https://www.ukinbound.org/resources/list-of-mp-twitter-accounts/</p> <p>&nbsp;</p> <p>If you use this dataset for an academic work, please reference the following paper:</p> <pre>@article{tacchi2022signed, title={Signed ego network model and its application to Twitter}, author={Tacchi, Jack and Boldrini, Chiara and Passarella, Andrea and Conti, Marco}, journal={arXiv preprint arXiv:2206.15228}, year={2022} }</pre>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Moral Values of Twitter COVID-19 Vaccine Data

<p>This data is part of our accepted paper &quot;Learning to Adapt Domain Shifts of Moral Values via Instance Weighting&quot; at&nbsp;the 33rd ACM Conference on Hypertext and Social Media (HT &rsquo;22). We annotate moral values of COVID-19 vaccine-related tweets.&nbsp;</p>

opencc-by-3.0-usApr 2022View details →
zenodo40/100

Historikertage auf Twitter (2012-2018). Datenreport und Datenset

<p>This data report contains the annotated figures, statistics and visualisations of the project &quot;Die twitternde Zunft. Historikertage auf Twitter (2012-2018)&quot; by Mareike K&ouml;nig and Paul Ramisch. In addition, the methodological approach to corpus creation, data cleaning, coding, network and text analysis as well as the legal and ethical considerations of the project are described.</p> <p>The datasheets contain the dehydrated and annotated tweet ids that were used for our study. With the Twitter API this can be used to hydrate and restore the whole corpus, apart from deleted tweets. There are two versions of the CSV file, one with clean id values, the other where the id values are prepended with an &ldquo;x&rdquo;. This prevents certain tools from using scientific notation for the ids and breaking them, with the R library rtweet function read_twitter_csv() this is automatically resolved on import.</p> <p>The files contain the following data:</p> <ul> <li>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; status_id: The Twitter status id of the tweet</li> <li>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; corpus_user_id: A corpus specific id for each user within the corpus (not the Twitter user id)</li> <li>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; hauptkategorie_1: Primary category</li> <li>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; hauptkategorie_2: Primary category 2</li> <li>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Gender: Gender of the user</li> <li>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Nebenkategorie: Secondary category</li> </ul> <p>Furthermore, the following boolean variables describe what sub corpus each tweet is in, the main corpus per year that contains of both data sources (TAGS and API) and the yearly sub corpora divided by their data source (TAGS: orig_, API: api_):</p> <p>You can find the code on R on GitHub: <a href="https://github.com/dhiparis/historikertag-twitter">https://github.com/dhiparis/historikertag-twitter</a>.</p>

opencc-by-4.0Mar 2022View details →
dryad40/100

Twitter data reveal six distinct environmental personas

<p>Effective digital environmental communication is integral to galvanizing public support for conservation in the age of social media. Environmental advocates require messaging strategies suited to social media platforms, including ways to identify, target, and mobilize distinct audiences. Here, we provide – to the best of our knowledge – the first systematic characterization of environmental personas on social media. Beginning with 1 million environmental nongovernmental organization (NGO) followers on Twitter, of which 500,000 users met data quality criteria, we identified six personas that differ in their expression of 21 environmental issues. General consistency in the proportional composition of personas was detected across 14 countries with sufficiently large samples. Within the US, although the six personas varied in their mean political ideology, we did not observe that the personas split along political party lines. Our results pave the way for environmental advocates – including NGOs, public agencies, and researchers – to use audience segmentation methods like the one discussed here to target and tailor messages to distinct constituencies at speed and scale. This repository contains several tabular files that can be used to query user data from Twitter or reproduce the main results in the main text of the article.</p>

opencc-zeroMay 2022View details →
zenodo40/100

Radical Right On Twitter (ROT)

<p>We collected the Radical Right On Twitter dataset (ROT7) to advance research into radical right activity online. The resource addresses a lack of data in this field, particularly data that relates to the activity of radical right actors. The dataset was funded without commercial support.</p> <p>ROT follows six months of Twitter activity (8th July 2020 to 9th January 2021) from 35 radical right actors. We follow the advice given by Williams, Burnap, and Sloan (2017) for publishing Twitter data on sensitive topics. ROT includes:</p> <p>It contains:</p> <ol> <li>&nbsp;Actors&#39; content: all content produced by the actors, including posts (n = 22,131), replies (n = 19,947), quotes (n = 11,314), and retweets (n = 37,283).</li> <li>Actors&#39; profiles: Twitter profile information for all 35 actors.</li> <li>Actors&#39; followers: a list of each actor&#39;s followers, collected each day (combined n = 6,592,056).</li> <li>Actors&#39; friends: a list of each actor&#39;s friends, collected each day (combined n = 262,856).</li> <li>Direct engagement: all tweets which engage with actors, including replies, quotes and retweets (n = 31,443,828).</li> <li>Engagers&rsquo; followers: List of followers of every user who replied, quoted or retweeted actors&#39; content. We only collected users&#39; list of followers once, even if they engaged with the actors multiple times during the period studied.</li> <li>Other engagement: all other tweets collected through Twitter API that mentions an actor (n = 10,939,868).&nbsp;</li> </ol>

opencc-by-4.0May 2022View details →
zenodo40/100

News headlines of BBC articles published by @BBCBreaking twitter account

<p>The dataset consists of a list of news articles headlines retrieved from tweets published by @BBCBreaking profile in specific years (2012, 2015, 2017, 2019 and 2022).</p> <p>The dataset is in&nbsp;<code>.csv</code>&nbsp;format and is organised as follows:</p> <ul> <li>Columns: <ul> <li>ID (tweet ID)</li> <li>created_at (tweet publication&#39;s date)</li> <li>url (url of the news article attached to the tweet)</li> <li>Titles (news headline)</li> </ul> </li> <li>Rows: Each row contains a single news article headline sorted by date of publication (created_at). Total number of entries: 7213.</li> </ul> <p>For more details about data collection refer to <a href="https://github.com/caiocmello/news-mood">Github</a>.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Política, Twitter y Covid-19. La percepción de los turistas extranjeros durante la desescalada en España en la primavera de 2021

<p>This file contains the job database described below.&nbsp;The traditional journalism of the written press of the day and the linear generalist television news have lost their monopoly as generators of public opinion. The communicative and social relevance of social networks is indisputable, even more so in circumstances such as those provoked by the Covid-19 election period for the Community of Madrid. We focus on two issues of great relevance in networks in this period: the image of tourism and policy makers on Twitter and in Spain between March and May 2021. The results of the analyses carried out with computational and corpus tools reveal a very heterogeneous constellation of participants, as well as that the pandemic and its management was, in general, used as an electoral weapon through the arrival of tourism in the middle of the pandemic in the capital. The analyses of key words and their associations show a vision of tourism with a clear electoral intention, where all the general interest was subordinated to the battle for Madrid.</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record