Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
169
datasets available to search
ShareScore release 0.7.1
Dataset results
169 results for “tweet”
Tweets from the European Patent Office account (@epoorg): April 2009-July 2022
<p>19566 tweets released by th the European Patent Office account (@epoorg): April 2009-July 2022.</p> <p>Fields: id<gx:category>, author_id<gx:category>, author_name<gx:category>, author_handler<gx:category>, author_avatar<gx:url> ,user_created_at<gx:date>, user_description<gx:text>, user_favourites_count<gx:number>, user_followers_count<gx:number>, user_following_count<gx:number>, user_listed_count<gx:number>, user_tweets_count<gx:number>, user_verified<gx:boolean>, user_location<gx:text>, lang<gx:category>, type<gx:category>,text<gx:text>, date<gx:date>, mention_ids<gx:list[category]>, mention_names<gx:list[category]>, retweets<gx:number>, favorites<gx:number>, replies<gx:number>, quotes<gx:number>, links<gx:list[url]>, links_first<gx:url>, image_links<gx:list[url]>, image_links_first<gx:url>, rp_user_id<gx:category>, rp_user_name<gx:category>, location<gx:text>, tweet_link<gx:url>, source<gx:text>, search<gx:category></p>
The #BTW17 Twitter Dataset - Recorded Tweets of the Federal Election Campaigns of 2017 for the 19th German Bundestag
<p>The German Bundestag elections are the most important democratic elections of Germany. This dataset comprises Twitter interactions related with German politicians of the most important political parties over several months in the (pre-)phase of the German election campaigns in 2017. The Twitter accounts of 364 politicians (that is approximately half of the German parliament, the German Bundestag) were followed for almost half a year. The collected data comprise of about 10 GB of Twitter raw data generated by more than 120.000 active Twitter users generating more than 1.200.000 tweets during the pre- and hot-phase of the election campaigns for the 19th German Bundestag. <br> The dataset can be used to study how political parties, their followers and supporters make use of social media channels like Twitter in the context of political election campaigns and what kind of content is shared.</p> <p>The following files contain relevant context information:</p> <ul> <li><strong>crawled-pages.json</strong> contains the URLs of the official party faction websites of the 18th German Bundestag that were crawled to identify the Twitter screennames of German politicians of all Bundestag factions. Because the <em>Alternative für Deutschland (AfD)</em> and the <em>Freie Demokratische Partei (FDP)</em> were not part of the 18th German Bundestag (but it was likely that they will enter the 19th German Bundestag) other official websites were selected to crawl for relevant and representative politicians for these both parties (in case of the <em>AfD</em> this was the website of the directorate of the <em>AfD</em> federal party and the list of members of the European Parliament, in case of the <em>FDP</em> this was the website of the executive committee of the <em>FDP</em> federal party of Germany).</li> <li><strong>followed-accounts.json</strong> contains the (manually checked and edited) crawling result of 327 Twitter screennames of politicians that have been observed via the Twitter streaming API to collect this dataset.</li> </ul>
Tweet_Eleições_2022: Um dataset de tweets durante as eleições presidenciais brasileiras de 2022
<p>O dataset foi criado a partir da API da plataforma Twitter durante o período das eleições de 2022 no Brasil, com o intuito de capturar tweets relacionados às eleições e aos temas políticos relevantes. </p> <p>Identificamos e selecionamos as notícias e acontecimentos relacionados às eleições de 2022. Isso envolveu monitorar diversos canais de informação, como sites de notícias, portais on-line e jornais, a fim de identificar temas relevantes e hashtags em destaque. A seleção das notícias foi realizada com base no interesse e avaliação dos autores sobre a sua repercussão na mídia oficial e o quanto se associava direta ou indiretamente a personalidades ou a ideologias político-partidárias presentes nas eleições.</p> <p>Na sequência, foram selecionadas palavras-chave e hashtags pertinentes aos eventos escolhidos. Esses termos foram utilizados para configurar a query de extração, especificando o período de interesse, o limite máximo de tweets a serem recuperados e diversos campos do tweet foram coletados para fins da pesquisa para qual o dataset foi inicialmente gerado.</p> <p> </p>
Blood Donation Japanese Tweets - Dehydrated versions
<p>These datasets contains tweets collected through Twitter API with the previously available Research Account, during May 2023. Data was collected by using the “(献血 OR kenketsu OR けんけつ OR ラブラッド OR LoveBlood OR #献血 OR #けんけつ OR #kenketsu OR #ラブラッド OR #LoveBlood) lang:ja” search string, which containts the japanese words related to Blood Donation in japanese, and the name of the official blood donation application from Japan, LoveBlood.</p> <p>The Twitter_labeled_dataset contains tweets from the year 2022 randomly selected to prepare for manual labeling, while the other two datasets contain the tweets collected from October 2022 to April 2023 for automatized classification. Dataset_unlabeled_2023_full references to the raw collected tweets, while Data_classified_full has the final group of tweets after preprocessing and additional filtering.</p> <p>Considering the privacy recommendations when using Twitter data, all datasets has been "dehydrated," meaning they only contains the tweet IDs and not the associated content. To use this data, you will need to reassociate the IDs with their corresponding data (rehydration).</p>
Tweeting after terrorism: A content analysis of Twitter crisis communication following the 2016 Brussels Bombings, Dataset
<p>This study attempted to further the literature by showing how crisis communication and social media share a common foundation in strategic communication management. A content analysis, along with insights from a social network analysis, was used to better understand not only how information diffused across a governmentally-centered Twitter network following the 2016 Brussels bombings, but how specific types of strategic digital content performed. A sample of 1,187 tweets suggested that emotionally-neutral tweets from government agencies functionally diffused better across the social network, while more sentimental tweets from politicians performed better in a reputation management context. Such findings contribute to the literature on communication, political science, and management scholarship by conceptually and empirically showing how interdisciplinary research can serve as a strategic bridge between a multitude of actors interested in crisis communication research and practice.</p>
Italian Tweet Embeddings Used For Emoji Prediction
<p>This dataset contains 100d word embeddings trained on 48M Italian tweets using fastText and employed by our team to predict emojis during ITAmoji competition of EVALITA 2018 Evaluation Campaign.</p>
Dataset of tweets, used to detect hazardous events at the Baths of Diocletian site in Rome
<p>This dataset is composed by 276865 tweets, extracted from the Twitter stream, in the period from May 2018 to May 2019. Each element is characterized by the ID, that permits to retrieve the tweet from the stream, the text content, the GPS information (if it is provided or not), the localization, and the time. The last attribute of the elements of the dataset is the label associated to the tweet after the process of event detection. If the tweet has been identified as containing useful information of an hazardous event at the Baths of Diocletian site in Rome, the last attribute (Generated Useful Info at BOD) is "yes", otherwise its value is "no".</p>
Tweet IDs for the #brexit tweet dataset collected for the Helsinki Digital Humanities Hackathon 2019
<p>This dataset contains lists of tweet ids for the tweets used as material by the "Brexit in Transna­tional So­cial Me­dia" group in the <a href="http://heldig.fi/dhh19/">Helsinki Digital Humanities Hackathon 2019</a>.</p> <p>Due to restrictions in Twitter's terms of service, the full tweet dataset cannot be made public. However, Twitter allows the publication of tweet ids, from which the dataset can be reconstituted, <em>with the exception of deleted tweets</em>.</p> <p>The dataset was gathered as follows: Between 2019-01-22T09:19Z and 2019-04-15T14:42Z, a <a href="https://github.com/DocNow/twarc">Twarc</a> version 1.6.1 script was called hourly to retrieve and archive tweets from Twitter matching the #brexit hashtag using the Twitter search API. All 5,547,585 tweet IDs returned by this run are listed in the file <code>original_ids.txt.gz</code>.</p> <p>However, upon further inspection, problems were identified in the archiving. For an unidentified reason, gathering did not occur between 2019-02-13T06:17Z and 2019-02-26T09:24Z. In addition, the Twarc script had been run without the <code>--extended</code> argument, so the tweet data contained only truncated contents for many tweets.</p> <p>Due to this, a decision was taken to rehydrate a new dataset of tweets falling between 2019-02-26T09:24Z and 2019-04-15T14:42Z. Of the original 4,197,059 tweets gathered for this time period (listed in <code>continuous_ids.txt</code>), 3,941,653 could be rehydrated (i.e., they had not been deleted in the interim). The ids of tweets in this dataset are listed in the file <code>continuous_rehydrated_ids.txt</code>.</p> <p>Finally, out of these rehydrated tweets, a subset was derived that filtered out all tweets that were pure retweets. This subset consisted of 1,104,514 tweets, whose ids are listed in the file <code>continuous_rehydrated_no_retweets_ids.txt</code>.</p>
SET DE TWEETS FRANCOPHONES RELATIF A LA CRISE DE COVID-19 A DES FINS DE RECHERCHE
<p>Contexte</p> <p>Ce dataset est mis à disposition dans le cadre du projet MESCOV « les media sociaux lors de la crise Covid-19 » financé par le Comité analyse, recherche et expertise (CARE) du Ministère de l’Education Supérieure, de la Recherche et de l’Innovation. Le projet MESCOV traite des aspects création et circulation de l’information sur les media sociaux lors de la crise COVID-19, des initiatives citoyennes qui y ont émergé et des pratiques des professionnels de la gestion de crise associées (notamment Service d’Incendie et de Secours et Préfecture).</p> <p>C’est un projet pluridisciplinaire qui mobilise à la fois les Sciences de l’Informatique et de la Donnée pour le module base de données et algorithme d’apprentissage automatique et les Sciences Humaines et Sociales pour la partie documentation des mécanismes de création, de circulation et de vérification de l’information sur les media sociaux, l’émergence d’initiatives citoyennes et l’utilisation des médias sociaux par les institutionnels (Camozzi, et al., 2020, à paraître).</p> <p>Acquisition des données et constitution du jeu de tweets</p> <p>Ce jeu de tweets a été généré à partir du dataset proposé par Banda et al. (2020) récolté en temps-réel. Ce dataset a été constitué à partir des mots-clés suivants : <em>COVD19, CoronavirusPandemic, COVID-19, 2019nCoV, CoronaOutbreak,coronavirus , WuhanVirus, covid19, coronaviruspandemic, covid-19, 2019ncov, coronaoutbreak, wuhanvirus. </em></p> <p>Les tweets utilisés sont ceux publiés entre le 22 mars et le 24 juin 2020.</p> <p>Le jeu de données que nous proposons a été hydraté en utilisant l’outil Twarc, et seuls les tweets en langue française ont été conservés. Il comprend 2.950.157 tweets au 15 juillet 2020, date de sa création.</p> <p>Conditions d’utilisation et citation du set de tweets</p> <p><strong>Si vous citez ou réutilisez ce dataset merci de mentionner Montarnal, A., Coche, J., Bubendorff, S. Thubert, N., Camozzi, M-L., & Rizza, C. (2020) « Set de tweets francophones relatif à la crise de covid-19 à des fins de recherche ».</strong></p> <p><em>Ce jeu de tweets a été créé sur la base légale de l’intérêt public. Il est fourni tel quel, suivant les</em></p> <p><em>règlements de Twitter.</em> Seuls les identifiants des tweets sont fournis au format csv. L’hydratation du dataset est possible, par exemple en utilisant le projet Python Twarc (<a href="https://github.com/DocNow/twarc">https://github.com/DocNow/twarc</a>).</p> <p>Il ne contient pas de données personnelles mais en réhydratant ce dataset, les chercheurs pourront retrouver les auteurs des tweets à partir de l’ID du tweet associé. Ce jeu de tweets ne peut être utilisé qu’à des fins non-commerciales et de recherche. Les chercheurs seront responsables de leur traitement.</p> <p>Références</p> <p>Banda J-M., <em>et al.</em>, “A large-scale COVID-19 Twitter chatter dataset for open scientific research -- an international collaboration,” <em>arXiv:2004.03688 [cs]</em>, Nov. 2020, Accessed: Nov. 24, 2020. [Online]. Available: <a href="http://arxiv.org/abs/2004.03688">http://arxiv.org/abs/2004.03688</a>.</p> <p>Camozzi M-L., Thubert N., Coche J., Bubendorff S., Montarnal A., & Rizza C. (2020) Les media sociaux lors de la crise sanitaire de Covid-19 : Circulation de l'information et initiatives citoyennes. i3 Working Papers Series, 16-SES-01</p>
Pre-processed Tweets by verified users, Elon Musk, Vitalik Buterin and CZ Binance
<p>This dataset comprises of tweets from verified users, Elon Musk, Vitalik Buterin and CZ Binance that have keywords "bitcoin", "cryptocurrency", "btc" and "crypto"</p>
Tweets acerca dos imaginários algorítmicos na venda de livros pela Amazon
<p>Esta planilha compõe a metodologia da pesquisa de mestrado intitulada: Imaginários Algorítmicos e do ciclo de sedução da Amazon: um estudo de caso sobre a mediação algorítmica e a venda de livros na plataforma. Com base em sua elaboração, analisamos 1.065 tweets coletados, durante um período de três meses (04/02/2022 a 04/05/2022), com o objetivo de mapear os sentidos produzidos pelos usuários, no contexto de interação com a plataforma, e os imaginários que decorrem deles.</p> <p>Para tanto, elaboramos quatro categorias de análise, seguindo o proposto pela Análise de Conteúdo de Bardin (2016). São elas: 1) mobilização do imaginário por meio dos sentidos; 2) mobilização do imaginário durante a relação de consumo com a <em>Amazon</em>; 3) reconhecimento pelo usuário das ações algorítmicas; 4) referente a lógica de atuação da <em>Amazon</em> na venda de livros.<br> <br> A tabela é composta por 14 colunas, separadas por letras do alfabeto. A) refere-se ao tweet ID; B) a identificação da conta do usuário do Twitter; C) o texto elaborado pelo usuário; D) Categoria 01; E) Categoria 02; F) Categoria 03; G) Categoria 04; H) Tipo de tweet (compõe os imaginários, perfis de vendas ou outros assuntos que divergem do foco da pesquisa); I) Tweets repetidos que também compõem os perfis de vendas e afiliados da Amazon; J) Retweets; K) replys (R = respostas e M= menções; L) data e hora dos posts; M) idioma; N) número de seguidores dos usuários. </p>
Coronavirus Twitter Data: A collection of COVID-19 tweets with automated annotations
<p>This dataset contains tweets related to COVID-19. The dataset contains Twitter ids, from which you can download the original data directly from Twitter. Additionally, we include the date, keywords related to COVID-19 and the inferred geolocation. Check detailed information at <a href="http://twitterdata.covid19dataresources.org/index">http://twitterdata.covid19dataresources.org/index</a>.</p>
Análisis de los Tweets con mayor Interacción de las cuentas del Gobierno de España de Educación y Cultura
<p>Selección de los 10 Tweets con mayor número de "me-gustas" y de Retweet (eliminados los duplicados) de las cuentas vinculadas al Ministerio de Cultura y Deporte y Ministerio de Educación y Formación Profesional, durante el periodo: marzo 2020 a marzo 2021.</p>
MR Lit Tweets
<p>Raw data on tweet times, content, impression, retweets etc from MR Lit</p> <p>zip file added after. contains all the separate files</p> <p>combined file added which is all the separate csvs together.</p> <p>maintains only columns Tweet id,Tweet permalink,Tweet text,time,impressions,engagements,engagement rate,retweets,replies,likes,user profile clicks,url clicks. added date column.</p>
IDs of Tweets containing data on Work Life balance before and after covid
<p>This dataset contains Tweet IDs that contain " "work-life balance" OR "worklife balance" OR "#work-lifebalance" OR "#worklifebalance". 1,768,628 tweets were downloaded based on this query. Specifically 915 402 „before covid“ and 853 226 „after covid“.<br> This dataset contains only Tweet IDs.</p>
[Dataset] Tweets about COVID-19 Brazilian PCI
<p>Installed in April 2021, the COVID-19 Parliamentary Commission of Inquiry (PCI) aimed to investigate omissions and irregularities committed by the federal government during the COVID pandemic in Brazil, which resulted in the death of more than 660,000 Brazilians and placed it among the countries with the most deaths caused by COVID-19.</p> <p>This dataset has 3,397,933 tweets, splitted in days and weeks, extracted over a period of 26 weeks. It contains textual data from tweets, data about users (@ and description), and data about interactions between users. It can be used to improve textual cleaning techniques, toxic speech detection, clustering, and even Social Network Analysis and social graph studies. Data format is parquet.</p> <p>This dataset is part of a [paper](https://doi. org/10.1145/3539637.3556992)[1], published by its author, which aimed to do a social network analysis related to the CPI topic, to investigate evidence of political polarization. The source codes and jupyter notebooks are available on <a href="https://github.com/lucasraniere/twitter-cpi-case-study">GitHub</a>.</p> <p>[1] Uniting Politics and Pandemic: a Social Network Analysis on the COVID Parliamentary Commission of Inquiry in Brazil. WebMedia 2022. Lucas Raniére J. Santos, Leandro B. Marinho, Caludio E. C. Campelo.</p>
Twitter historical dataset: March 21, 2006 (first tweet) to July 31, 2009 (3 years, 1.5 billion tweets)
<p><strong>Disclaimer: </strong>This dataset is distributed by Daniel Gayo-Avello, an associate professor at the Department of Computer Science in the University of Oviedo, for the sole purpose of non-commercial research and it just includes tweet ids.</p> <p>The dataset contains tweet IDs for <strong>all the published tweets</strong> (in any language) <strong>bettween March 21, 2006 and July 31, 2009</strong> thus comprising the first whole three years of Twitter from its creation, that is, about <strong>1.5 billion tweets</strong> (see file <em>Twitter-historical-20060321-20090731.zip</em>).</p> <p>It covers several defining issues in Twitter, such as the invention of hashtags, retweets and trending topics, and it includes tweets related to the 2008 US Presidential Elections, the first Obama’s inauguration speech or the 2009 Iran Election protests (one of the so-called Twitter Revolutions).</p> <p>Finally, it does contain tweets in many major languages (mainly English, Portuguese, Japanese, Spanish, German and French) so it should be possible–at least in theory–to analyze international events from different cultural perspectives.</p> <p>The dataset was completed in November 2016 and, therefore, the tweet IDs it contains were publicly available at that moment. This means that there could be tweets public during that period that do not appear in the dataset and also that a substantial part of tweets in the dataset has been deleted (or locked) since 2016.</p> <p>To make easier to understand the decay of tweet IDs in the dataset a number of representative samples (99% confidence level and 0.5 confidence interval) are provided.</p> <p>In general terms, <strong>85.5% ±0.5 of the historical tweets are available as of May 19, 2020</strong> (see file <em>Twitter-historical-20060321-20090731-sample.txt</em>). However, since the amount of tweets vary greatly throughout the period of three years covered in the dataset, additional representative samples are provided for 90-day intervals (see the file <em>90-day-samples.zip</em>).</p> <p>In that regard, the ratio of publicly available tweets (as of May 19, 2020) is as follows:</p> <ul> <li>March 21, 2006 to June 18, 2006: 88.4% ±0.5 (from 5,512 tweets).</li> <li>June 18, 2006 to September 16, 2006: 82.7% ±0.5 (from 14,820 tweets).</li> <li>September 16, 2006 to December 15, 2006: 85.7% ±0.5 (from 107,975 tweets).</li> <li>December 15, 2006 to March 15, 2007: 88.2% ±0.5 (from 852,463 tweets).</li> <li>March 15, 2007 to June 13, 2007: 89.6% ±0.5 (from 6,341,665 tweets).</li> <li>June 13, 2007 to September 11, 2007: 88.6% ±0.5 (from 11,171,090 tweets).</li> <li>September 11, 2007 to December 10, 2007: 87.9% ±0.5 (from 15,545,532 tweets).</li> <li>December 10, 2007 to March 9, 2008: 89.0% ±0.5 (from 23,164,663 tweets).</li> <li>March 9, 2008 to June 7, 2008: 66.5% ±0.5 (from 56,416,772 tweets; see below for more details on this).</li> <li>June 7, 2008 to September 5, 2008: 78.3% ±0.5 (from 62,868,189 tweets; see below for more details on this).</li> <li>September 5, 2008 to December 4, 2008: 87.3% ±0.5 (from 89,947,498 tweets).</li> <li>December 4, 2008 to March 4, 2009: 86.9% ±0.5 (from 169,762,425 tweets).</li> <li>March 4, 2009 to June 2, 2009: 86.4% ±0.5 (from 474,581,170 tweets).</li> <li>June 2, 2009 to July 31, 2009: 85.7% ±0.5 (from 589,116,341 tweets).</li> </ul> <p>The apparent drop in available tweets from March 9, 2008 to September 5, 2008 has an easy, although embarrassing, explanation.</p> <p>At the moment of cleaning all the data to publish this dataset there seemed to be a gap between April 1, 2008 to July 7, 2008 (actually, the data was not missing but in a different backup). Since tweet IDs are easy to regenerate for that Twitter era (source code is provided in <em>generate-ids.m</em>) I simply produced all those that were created between those two dates. All those tweets actually existed but a number of them were obviously private and not crawlable. For those regenerated IDs the actual ratio of public tweets (as of May 19, 2020) is 62.3% ±0.5.</p> <p>In other words, what you see in that period (April to July, 2008) is not actually a huge number of tweets having been deleted but the combination of deleted *and* non-public tweets (whose IDs should not be in the dataset for performance purposes when rehydrating the dataset).</p> <p>Additionally, given that not everybody will need the whole period of time the earliest tweet ID for each date is provided in the file <em>date-tweet-id.tsv</em>.</p> <p>For additional details regarding this dataset please see: <strong>Gayo-Avello, Daniel. "How I Stopped Worrying about the Twitter Archive at the Library of Congress and Learned to Build a Little One for Myself." <em>arXiv preprint arXiv:1611.08144</em> (2016)</strong>.</p> <p><strong>If you use this dataset in any way please cite that preprint </strong>(in addition to the dataset itself).</p> <p>If you need to contact me you can find me as <a href="https://twitter.com/pfcdgayo">@PFCdgayo</a> in Twitter.</p>
replies and retweets with comments (quotes) of the tweets 1261179829895995394 and 1263036802295902208
<p>The dataset contains (recursive) replies and retweets with comments, as well as their (recursive) replies of two "initial" tweets: 1261179829895995394 and 1263036802295902208.</p> <p>The code that was used is published here: <a href="https://github.com/millawell/reretweets">https://github.com/millawell/reretweets</a></p> <p>The script takes the following steps to obtain tweets and quotes:</p> <ol> <li>It downloads all direct replies (recursively) to the tweet</li> <li>It searches for the tweet id, to get quotes</li> <li>For each tweet, that quoted the original tweet, it again retrieves all replies.</li> </ol>
eNGO_tweets
<p>Content scraped from the Twitter timelines of ten leading eNGO accounts. Column headers are as per the get_timelines() function in the R package rtweet (https://rtweet.info/). An additional column of manually assigned biodiversity threat categories has been added, broadly following the IUCN classification scheme. See, for example: Maxwell, S. L. et al. Nature 536, 143–145 (2016).</p>
Tweet IDs using History related hashtags
<p>This repository contains IDs of tweets that are related to history and that were collected for the purpose of analyzing how history-related content is disseminated in online social networks. Our <a href="https://link.springer.com/article/10.1007/s00799-020-00296-2">IJDL paper</a> shows the analysis results. The preliminary version of the analysis report is available <a href="https://dl.acm.org/doi/10.1145/3197026.3197057">here</a>.</p> <p> </p> <p>We used the <a href="https://developer.twitter.com/en/docs/tweets/search/api-reference/get-search-tweets.html">Twitter official search API</a> provided by Twitter to collect tweets. Note that three kinds of tweets are typically found in Twitter: tweets, retweets and quote tweets. Tweet is an original text issued as a post by a Twitter user. A retweet is a copy of an original tweet for the purpose of propagating the tweet content to more users (i.e., one's followers). Finally, a quote tweet copies the content of another tweet and allows also to add new content. A quote tweet is sometimes called a retweet with a comment. In this work, we simply treat all quote tweets as original tweets since they include additional information/text. There were however only 1,877 (0.2%) tweets recognized as quote tweets in our dataset. </p> <p> </p> <p>To collect tweets that refer to the past or are related to collective memory of past events/entities, we performed hashtag based crawling together with bootstrapping procedure. <br> At the beginning, we gathered several <a href="http://blog.historians.org/2013/08/history-hashtags-exploring-a-visual-network-of-twitterstorians/">historical hashtags selected by experts</a> (e.g. <strong>#HistoryTeacher</strong>, <strong>#history</strong>, <strong>#WmnHist</strong>). <br> In addition, we prepared several hashtags that are commonly used when referring to the past: <strong>#onthisday</strong>, <strong>#thisdayinhistory</strong>, <strong>#throwbackthursday</strong>, <strong>#otd</strong>. We then collected tweets that contain these hashtags by using Twitter official search API.</p> <p> </p> <p>The collected tweets were issued from 8 March 2016 to 2 July 2018. <br> Bootstrapping allowed us to search for other hashtags frequently used with the seed hashtags. The tweets tagged by such hashtags were then included into the seed set after the manual inspection of all the discovered hashtags as of their relation to the history, and filtering ones that are unrelated. <br> In total, we gathered 147 history-related hashtags which allowed us to collect 2,370,252 tweet IDs pointing to 882,977 tweets and 1,487,275 re-tweets.</p> <p> </p> <p>Related papers:</p> <ol> <li>Yasunobu Sumikawa, Adam Jatowt, and Marten During, <strong>"Digital History meets Microblogging: Analyzing Collective Memories in Twitter"</strong>, In Proceedings of the 18th ACM/IEEE-CS Joint Conference on Digital Libraries, JCDL'18, IEEE/ACM, pp. 213 -- 222, 2018. [<a href="https://dl.acm.org/doi/10.1145/3197026.3197057">paper</a>]</li> <li>Yasunobu Sumikawa and Adam Jatowt, <strong>"Analyzing History-related Posts in Twitter"</strong>, International Journal on Digital Libraries, Springer, 2020. https://doi.org/10.1007/s00799-020-00296-2 [<a href="https://link.springer.com/article/10.1007/s00799-020-00296-2">paper</a>]</li> <li>Yasunobu Sumikawa and Adam Jatowt, <strong>"Annotated Dataset of History-related Tweets"</strong>, Data in Brief, Vol. 38, pp. 107344, Elsevier, 2021. [<a href="https://www.sciencedirect.com/science/article/pii/S2352340921006284">paper</a>][<a href="https://zenodo.org/record/4657223">annotated dataset</a>]</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.