Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12
datasets available to search
ShareScore release 0.9.0
Dataset results
12 results for “hashtags”
#IndonesiaHumanRightsSOS Twitter Hashtag Tweets Dataset
<p>Dataset ini merupakan hasil dari scraping pada media sosial twitter dengan menggunakan aplikasi twint yang ditujukan pada hashtag #IndonesiaHumanRightsSOS. Scraping data dilakukan untuk cuitan yang dibuat dari tanggal 18 Desember 2020 10:59 AM s/d 19 Desember 2020 23:18 PM.</p> <p>Pada dataset mengandung 106.903 Row data dengan informasi terkait: User ID, Username, Twitter Name,Tweets, dsb.</p> <p>Selain itu dilampirkan juga contoh data yang telah dianalisis berupa wordcloud,username cloud, 100 most used word & most active username.</p> <p>-</p> <p>This dataset is the result of scraping on social media twitter using the twint application aimed at the hashtag #IndonesiaHumanRightsSOS. Data scraping is done for tweets made from December 18 2020 10:59 AM to December 19 2020 23:18 PM.</p> <p>The dataset contains 106,903 rows of data with related information: User ID, Username, Twitter Name, Tweets, etc.</p> <p>Also there is an example of the data that has been analyzed in the form of wordcloud, username cloud, 100 most used words & most active username.</p>
Twitter hashtags time series used in the paper "Universality, criticality and complexity of information propagation in social media"
<pre>These files contain the time series and the associated hashtags we obtained by sampling Twitter for our paper "Universality, criticality and complexity of information propagation on social media". The analysis is reported in <a href="https://arxiv.org/abs/2109.00116">https://www.nature.com/articles/s41467-022-28964-8</a> Please acknowledge the use of these data by citing the paper above. ################################# ################################# DATA ORGANIZATION We created a single zip file with all the time series and a single zip file with all the hashtags. There is a one-to-one correspondence between lines in the two files. ################################# ################################# FILES CONTENT As stated, here is a one-to-one correspondence between lines in the time series file and lines in the hashtags file, i.e., the hashtag stored in line X is the hashtag of the time series stored in line X. Time series are stored as follows: Ka t1 t2 t3 \n Kb t1 t2 t3 t4 t5 \n . . . Kn t1 t2 \n where: Ka, Kb,..., Kn is an integer specifying the number of events that compose the time series a, b,..., n respectively. In the example above we would have Ka=3, Kb=5, Kn=2. t1 t2 ... is the time series, i.e., a sequence of chronologically ordered interevent times. The last interevent time, in our implementation, represents the distance between the end of the temporal window and the last event time. It thus does not represent an event. As stated in the Supplemental Material of our paper, the temporal window ranges from 2019, October 1st to 2019, November 30th. </pre>
Dataset of Mastodon toots using the hashtag #Fairdata
<p>This dataset provides mastodon toots that are using the hashtag FAIR-Data. This dataset is supposed to provide the foundation of further network analysis around the topic FAIR-Data. </p> <p><strong>Data Collection</strong></p> <p>Data was harvested using the Mastodon API. For each Mastodon server listed in the dataset, the API was employed to retrieve posts tagged with "fairdata". Up to 5,000 posts were collected per request, utilizing the API's pagination feature to obtain all available posts for that hashtag. Each Mastodon server was queried separately, and the results were stored in distinct CSV files.</p> <p><strong>Potential Duplicates</strong></p> <p>Given that Mastodon operates as a federated network, a post made on one instance can be replicated across different instances. This implies that the same post might appear in the data from multiple Mastodon servers, likely accounting for the duplicates observed in the dataset.</p> <p><strong>Analysis of Federation Dynamics:</strong> Duplicates can reveal which content is shared between servers and which servers are most active in the federation. In this context, duplicates could provide valuable insights for the network analysis.</p> <p><strong>The columns are:</strong></p> <ol> <li><strong>id</strong> - A unique identifier for each post.</li> <li><strong>created_at</strong> - Timestamp indicating when the post was created.</li> <li><strong>content</strong> - The content or message of the post.</li> <li><strong>account</strong> - Detailed information about the account that created the post (ID, username, URL).</li> <li><strong>replies_count</strong> - Number of replies to the post.</li> <li><strong>reblogs_count</strong> - Number of reblogs or shares of the post.</li> <li><strong>favourites_count</strong> - Number of favorites or likes the post has received.</li> <li><strong>language</strong> - Language in which the post is written.</li> <li><strong>mentions</strong> - Any mentions in the post (for example, other users).</li> <li><strong>tags</strong> - Tags associated with the post.</li> <li><strong>emojis</strong> - Emojis used in the post.</li> <li><strong>category</strong> - Category of the post.</li> </ol>
Hashtag adoption time series
<p>The dataset released here has been used in our paper "#Bigbirds Never Die: Understanding Social Dynamics of Emergent Hashtags." [link to the paper](https://www.aaai.org/ocs/index.php/ICWSM/ICWSM13/paper/view/6083/6376)</p> <p>In this study, we examine the growth, survival, and context of over 250 novel hashtags during the 2012 U.S. presidential debates. Our analysis reveals the trajectories of hashtag use fall into two distinct classes: "winners" that emerge more quickly and are sustained for longer periods of time than other "also-rans" hashtags. Statistical analyses of the growth and persistence of hashtags reveal novel relationships between the hashtags' contextual features and the relative success of hashtags. This is the first study on the lifecycle of hashtag adoption and use in response to purely exogenous shocks, which has implications for understanding social influence and collective action in social media more generally.</p> <p>The dataset was the hashtag adoption time sequences during the four debates. They are used to create **Figure 1** in the paper (Cumulative tweet volume of hashtags over time, starting from each debate). </p> <p>Each csv file has three columns:<br> time (in UTC), y (the minute-by-minute cumulative tweet count), and tag (the hashtag name).</p> <p>The onset of the four debates are:<br> Debate 1: 2012-10-04 00:00:00 UTC<br> Debate 2: 2012-10-12 00:00:00 UTC<br> Debate 3: 2012-10-17 00:00:00 UTC<br> Debate 4: 2012-10-23 00:00:00 UTC</p> <p>The raw tweet data have been released via the ICWSM Data Sharing Service. See: <br> [http://www.icwsm.org/2013/datasets/datasets/](http://www.icwsm.org/2013/datasets/datasets/)</p> <p> </p> <p><strong>Publication</strong><br> If you make use of these data sets and code, please cite:</p> <p>Lin, Y.-R., Margolin, D., Keegan, B., Baronchelli, A., Lazer, D. (2013). #Bigbirds Never Die: Understanding Social Dynamics of Emergent Hashtags. In Proceedings of the 7th International AAAI Conference on Weblogs and Social Media (ICWSM 2013) </p>
Dataset used in the paper: "Scaling laws and dynamics of hashtags on Twitter"
<p>This dataset was used in the manuscript "Scaling laws and dynamics of hashtags on Twitter"..</p> <p>The Twitter data was obtained from a sample of 10% of all public tweets, provided by the Twitter streaming application programming interface. We extracted the hashtags from each tweet and counted how many times they were used in different time intervals. Time intervals of three different lengths were used: days, hours, and minutes. The tweets were published between November 1st 2015 and November 30th 2016, but not all time intervals between these dates are available.</p> <p>The <strong>four files</strong> in this dataset correspond each to one folder (collected using tar). Each folder contains compressed .csv files (compressed using gzip). The content of the .csv files in each folder are:</p> <p><em><strong>hashtags_frequency_day.tar</strong></em><br> Counts of hashtags in each day. The name of each file in the folder indicates the date (GMT). The entries in each file are the hashtag and the count in the interval.<br> <br> <em><strong>hashtags_frequency_hour.tar</strong></em><br> Counts of hashtags in each hour. The name of each file in the folder indicates the date (GMT). The entries in each file are the hashtag and the count in the interval.<br> <br> <em><strong>hashtags_frequency_minutes.tar</strong></em><br> Counts of hashtags in each minute. The name of each file in the folder indicates the date (GMT, only a fraction of all days is available). The entries in each file are the hashtag and the count in the interval.<br> <br> <em><strong>number_of_tweets.tar</strong></em><br> Counts of the number of tweets in each minute. The name of each file in the folder indicates the day. The entries in each file are the minute in the day (GMT) and count of tweets in our dataset.</p>
Tweet IDs using History related hashtags
<p>This repository contains IDs of tweets that are related to history and that were collected for the purpose of analyzing how history-related content is disseminated in online social networks. Our <a href="https://link.springer.com/article/10.1007/s00799-020-00296-2">IJDL paper</a> shows the analysis results. The preliminary version of the analysis report is available <a href="https://dl.acm.org/doi/10.1145/3197026.3197057">here</a>.</p> <p> </p> <p>We used the <a href="https://developer.twitter.com/en/docs/tweets/search/api-reference/get-search-tweets.html">Twitter official search API</a> provided by Twitter to collect tweets. Note that three kinds of tweets are typically found in Twitter: tweets, retweets and quote tweets. Tweet is an original text issued as a post by a Twitter user. A retweet is a copy of an original tweet for the purpose of propagating the tweet content to more users (i.e., one's followers). Finally, a quote tweet copies the content of another tweet and allows also to add new content. A quote tweet is sometimes called a retweet with a comment. In this work, we simply treat all quote tweets as original tweets since they include additional information/text. There were however only 1,877 (0.2%) tweets recognized as quote tweets in our dataset. </p> <p> </p> <p>To collect tweets that refer to the past or are related to collective memory of past events/entities, we performed hashtag based crawling together with bootstrapping procedure. <br> At the beginning, we gathered several <a href="http://blog.historians.org/2013/08/history-hashtags-exploring-a-visual-network-of-twitterstorians/">historical hashtags selected by experts</a> (e.g. <strong>#HistoryTeacher</strong>, <strong>#history</strong>, <strong>#WmnHist</strong>). <br> In addition, we prepared several hashtags that are commonly used when referring to the past: <strong>#onthisday</strong>, <strong>#thisdayinhistory</strong>, <strong>#throwbackthursday</strong>, <strong>#otd</strong>. We then collected tweets that contain these hashtags by using Twitter official search API.</p> <p> </p> <p>The collected tweets were issued from 8 March 2016 to 2 July 2018. <br> Bootstrapping allowed us to search for other hashtags frequently used with the seed hashtags. The tweets tagged by such hashtags were then included into the seed set after the manual inspection of all the discovered hashtags as of their relation to the history, and filtering ones that are unrelated. <br> In total, we gathered 147 history-related hashtags which allowed us to collect 2,370,252 tweet IDs pointing to 882,977 tweets and 1,487,275 re-tweets.</p> <p> </p> <p>Related papers:</p> <ol> <li>Yasunobu Sumikawa, Adam Jatowt, and Marten During, <strong>"Digital History meets Microblogging: Analyzing Collective Memories in Twitter"</strong>, In Proceedings of the 18th ACM/IEEE-CS Joint Conference on Digital Libraries, JCDL'18, IEEE/ACM, pp. 213 -- 222, 2018. [<a href="https://dl.acm.org/doi/10.1145/3197026.3197057">paper</a>]</li> <li>Yasunobu Sumikawa and Adam Jatowt, <strong>"Analyzing History-related Posts in Twitter"</strong>, International Journal on Digital Libraries, Springer, 2020. https://doi.org/10.1007/s00799-020-00296-2 [<a href="https://link.springer.com/article/10.1007/s00799-020-00296-2">paper</a>]</li> <li>Yasunobu Sumikawa and Adam Jatowt, <strong>"Annotated Dataset of History-related Tweets"</strong>, Data in Brief, Vol. 38, pp. 107344, Elsevier, 2021. [<a href="https://www.sciencedirect.com/science/article/pii/S2352340921006284">paper</a>][<a href="https://zenodo.org/record/4657223">annotated dataset</a>]</li> </ol>
IDs of Tweets from 1/1/2019-31/7/2022 that contain the hashtag #burnout.
<p>IDs of Tweets from 1/1/2019-31/7/2022 that contain the hashtag #burnout.</p>
Hashtag HPV: HPV Vaccine Twitter Education Program
ClinicalTrials.gov study NCT05204030. IPD Sharing: NO. Countries: 1. Publications: 3.
Dataset of Tweets for the paper: "Hashtag activism on Twitter: The effects of who, what, when and how a tweet is sent for promoting citizens' engagement with climate change"
Open the record for dataset details and reuse information.
Precision targeting of beta-catenin induces tumor reprogramming and immunity in hepatocellular cancers [Single-cell RNA-seq hashtag immune]
GEO Series GSE270974. Mus musculus. 2 samples. Type: Expression profiling by high throughput sequencing.
Comprehensive Collection of English, German, Russian and Ukrainian Tweets Containing the Word or Hashtag Ukraine During the Russian Invasion, February 2022 until May 2023
<p>Comprehensive dataset of Tweets containing the keyword 'ukraine' (in German, Russian and Ukrainian) as well as '#ukraine' (in English) since the Russian Ukraine Invasion in February 2022. The user handle column has been excluded to protect deleted accounts that have not been retweeted or replied to. Tweets have been collected via the Academic API using the Search endpoint in four languages:</p> <table> <tbody> <tr> <td><strong>Language</strong></td> <td><strong>Query</strong></td> <td><strong> Name in Dataset (column: event) </strong></td> <td><strong>Number of Tweets<br></strong></td> </tr> <tr> <td>English</td> <td>#ukraine AND lang='en' </td> <td>ukraine-en-hashtag</td> <td>45.8 million</td> </tr> <tr> <td>German</td> <td>ukraine AND lang='de' </td> <td>ukraine</td> <td>19.6 million</td> </tr> <tr> <td>Russian</td> <td>Украина AND lang:ru </td> <td>ukraine-ru</td> <td>5.02 million</td> </tr> <tr> <td>Ukrainian</td> <td>Україна AND lang:uk </td> <td>ukraine-uk</td> <td>4.1 million</td> </tr> </tbody> </table> <p><strong>Collection dates</strong></p> <p>Details on collection dates per Tweet (e.g. to compare with creation dates) as well as the IDs of Tweets for consistency checks can be found here: <a href="https://github.com/Leibniz-HBI/ukraine_twitter_data">https://github.com/Leibniz-HBI/ukraine_twitter_data</a> (<a href="https://doi.org/10.17605/OSF.IO/RTQXN">https://doi.org/10.17605/OSF.IO/RTQXN</a>)</p> <p><strong>File Naming Scheme</strong></p> <p>To enable downloads of selected timeframes and languages, the files are named by language, start and end date of the tweet creation timestamp.</p> <p><strong>Columns</strong></p> <p>The following columns are available:</p> <ul> <li>event: tag for query and language used for the query</li> <li>id: Tweet ID</li> <li>inserted_at: collection date</li> <li>last_updated_at: last update date (relevant for metrics such as follower count)</li> <li>text: Tweet text</li> <li>lang: language as determined by Twitter</li> <li>created_at: creation date of the Tweet</li> <li>conversation_id: Tweets with the same ID are part of the same reply tree to a tweet (provided by Twitter)</li> <li>author_follower_count: follower count of the Tweet's author account at the creation or last update time of the tweet</li> <li>replied_to: account the Tweet replies to</li> <li>replied_to_follower_count: follower count of the Tweet's replied to account at the creation or last update time of the tweet</li> <li>quoted: if quote tweet, ID of quoted tweet</li> <li>quoted_follower_count: analog to replied_to_follower_count</li> <li>retweeted: analog to quoted</li> <li>retweeted_follower_count: analog to replied_to_follower_count</li> <li>hashtags: hashtags of the Tweet</li> <li>urls: URLs in the tweet, shortened/unshortened, including links to media</li> <li>place_id: alphanumeric place ID provided by the Twitter API, mostly empty</li> </ul>
Weibo_hashtags_thematic_analysis_COVID-19
<p>This dataset is affiliated to the journal article: Xi, W., Xu, W., Zhang, X., Ayalon, L. A Thematic Analysis of Weibo Topics (Chinese Twitter Hashtags) Regarding Older Adults During the COVID-19 Outbreak, <em>The Journals of Gerontology: Series B</em>, Volume 76, Issue 7, September 2021, Pages e306–e312, <a href="https://doi.org/10.1093/geronb/gbaa148">https://doi.org/10.1093/geronb/gbaa148</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.