Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

307

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

307 results for “Twitter”

Learn how ShareScore rates datasets ↗
zenodo44/100

Twitter Poll: Is #OpenScience an essential rsrch skill Grad Schools should train in prep for #REF2020

<p>The Twitter Poll &quot;Is #OpenScience an essential rsrch skill Grad Schools should train in prep for #REF2020&quot; was run online in support of Horizon 2020 Project HEIRRI (Higher Education Insititutions &amp; Responsible Research &amp; Innovation) 1st Conference, 18 March 2016.</p> <p>The poll attracted 123 voters, 12,343 impressions and 517 engagements (4,2% conversion).</p> <p><em><strong>Event website:</strong></em><br /> HEIRRI 1st Conference http://heirri.eu/1st-heirri-conference/</p> <p><em><strong>CODE for EMBEDDING TWITTER POLL: </strong></em></p> <p>&lt;blockquote class=&quot;twitter-tweet&quot; data-lang=&quot;en&quot;&gt;&lt;p lang=&quot;en&quot; dir=&quot;ltr&quot;&gt;Is &lt;a href=&quot;https://twitter.com/hashtag/OpenScience?src=hash&quot;&gt;#OpenScience&lt;/a&gt; an essential rsrch skill Grad Schools should train in prep for &lt;a href=&quot;https://twitter.com/hashtag/REF2020?src=hash&quot;&gt;#REF2020&lt;/a&gt; ? &lt;a href=&quot;https://twitter.com/hashtag/OpenSci4Doc?src=hash&quot;&gt;#OpenSci4Doc&lt;/a&gt; &lt;a href=&quot;https://twitter.com/HEIRRI_&quot;&gt;@HEIRRI_&lt;/a&gt;&lt;/p&gt;&amp;mdash; Foster Open Science (@fosterscience) &lt;a href=&quot;https://twitter.com/fosterscience/status/709650182800068608&quot;&gt;March 15, 2016&lt;/a&gt;&lt;/blockquote&gt;<br /> &lt;script async src=&quot;//platform.twitter.com/widgets.js&quot; charset=&quot;utf-8&quot;&gt;&lt;/script&gt;</p>

opencc-by-4.0Mar 2016View details →
zenodo44/100

Twitter Poll: Should #OpenScience practices be part of tenure & #REF2020 criteria

<p>The Twitter Poll &quot;Should #OpenScience practices be part of tenure &amp; #REF2020 criteria?&quot; was run online in support of Dutch EU Presidency Open Science Meeting, 4-5 April 2016.</p> <p>The poll run from 1-8 April 2016, attracting 122 voters, 11,336 impressions and 469 engagements (4.1% conversion rate).</p> <p>The 122 Twitter voters are in alignement with one of the key Call for Action recommendations on &quot;New assessment, reward and evaluation systems New systems that really deal with the core of knowledge creation and account for the impact of scientific research on science and society at large, including the economy, and incentivise citizen science&quot;.</p> <p><em><strong>Event website</strong></em>: http://english.eu2016.nl/documents/reports/2016/04/04/amsterdam-call-for-action-on-open-science</p> <p><em><strong>CODE to EMBED TWITTER POLL:</strong></em></p> <p>&lt;blockquote class=&quot;twitter-tweet&quot; data-cards=&quot;hidden&quot; data-lang=&quot;en&quot;&gt;&lt;p lang=&quot;en&quot; dir=&quot;ltr&quot;&gt;Should &lt;a href=&quot;https://twitter.com/hashtag/OpenScience?src=hash&quot;&gt;#OpenScience&lt;/a&gt; practices be part of tenure &amp;amp; &lt;a href=&quot;https://twitter.com/hashtag/REF2020?src=hash&quot;&gt;#REF2020&lt;/a&gt; criteria? &lt;a href=&quot;https://twitter.com/EU2016NL&quot;&gt;@EU2016NL&lt;/a&gt; &lt;a href=&quot;https://twitter.com/openscience&quot;&gt;@openscience&lt;/a&gt; &lt;a href=&quot;https://twitter.com/hashtag/phdchat?src=hash&quot;&gt;#phdchat&lt;/a&gt; &lt;a href=&quot;https://twitter.com/PhDForum&quot;&gt;@PhDForum&lt;/a&gt;&lt;/p&gt;&amp;mdash; Foster Open Science (@fosterscience) &lt;a href=&quot;https://twitter.com/fosterscience/status/715777247899205636&quot;&gt;April 1, 2016&lt;/a&gt;&lt;/blockquote&gt;<br /> &lt;script async src=&quot;//platform.twitter.com/widgets.js&quot; charset=&quot;utf-8&quot;&gt;&lt;/script&gt;</p>

opencc-by-4.0Apr 2016View details →
zenodo44/100

URLs from tweets for a 2014 sample of Twitter users and for a set of computer scientists

<p>The files in this dataset are used to analyse the tweeting behaviour of computer scientists on Twitter. They comprise</p> <ul> <li>a set of 989,529 tweet-URL pairs (<em>tweets_2014_researcher.tsv.bz2</em>) from 2014 from 6,271 users of the computer scientists sample in https://zenodo.org/record/12942 specified by time, tweet id, user id, and URL,</li> <li>a set of 300,053,850 tweet ids (<em>tweets_2014_sample.tsv.bz2</em>) from the 1% Twitter stream sample from 2014,</li> <li>a set of 671,304 tweet-URL pairs (<em>tweets_2014_sample_6271_users.tsv.bz2</em>) from the 1% Twitter stream sample from 2014 for 6,271 users specified by time, tweet id, user id, and URL,</li> <li>a set of the top 10,000 host names (<em>MAG_hosts_10000.tsv</em>) from the Microsoft Academic Graph data (http://blogs.msdn.com/b/msr_er/archive/2015/06/26/announcing-the-microsoft-academic-graph-let-the-research-begin.aspx), specified by rank, URL count, and host name, and</li> <li>a set of 340 host names of URL shortening services (<em>url_shortening_services.tsv</em>).</li> </ul>

opencc-by-sa-4.0Sep 2016View details →
zenodo44/100

Social sensing of urban land use based on analysis of Twitter users' mobility patterns

<p>A companion dataset for the paper "Social sensing of urban land use based on analysis of Twitter users' mobility patterns". This dataset contains five files and one dictionary depicting the preferential return of Twitter users to their key locations and the urban land use types at these locations. More details can be found in the README file. </p>

opencc-by-4.0May 2017View details →
zenodo44/100

URLs from tweets for a 2014 sample of Twitter users and for a set of computer scientists

<p>The files in this dataset are used to analyse the tweeting behaviour of computer scientists on Twitter. They comprise</p> <ul> <li>a set of 989,529 tweet-URL pairs (<em>tweets_2014_researcher.tsv.bz2</em>) from 2014 from 6,271 users of the computer scientists sample in https://zenodo.org/record/12942 specified by time, tweet id, user id, and URL,</li> <li>a set of 300,053,850 tweet ids (<em>tweets_2014_sample.tsv.bz2</em>) from the 1% Twitter stream sample from 2014,</li> <li>a set of 605,080 tweet-URL pairs (<em>tweets_2014_sample_6694_users.tsv.bz2</em>) from the 1% Twitter stream sample from 2014 for 6,694 users specified by time, tweet id, user id, and URL,</li> <li>a set of the top 10,000 host names (<em>MAG_hosts_10000.tsv</em>) from the Microsoft Academic Graph data (http://blogs.msdn.com/b/msr_er/archive/2015/06/26/announcing-the-microsoft-academic-graph-let-the-research-begin.aspx), specified by rank, URL count, and host name, and</li> <li>a set of 340 host names of URL shortening services (<em>url_shortening_services.tsv</em>).</li> </ul> <p>In addition, the following rankings (based on the odds ratio) of domains, hosts, and URLs that appear in both the researcher dataset and the sample are included:</p> <ul> <li><em>domains_by_odds_ratio.tsv.bz2</em> - a ranking of 61,860 domains,</li> <li><em>hosts_by_odds_ratio.tsv.bz2</em> - a ranking of 80,384 hosts,</li> <li><em>publisher_domains_by_odds_ratio.tsv.bz2</em> - a ranking of 924 publisher domains,</li> <li><em>publisher_urls_by_odds_ratio.tsv.bz2</em> - a ranking of 4,227 publisher URLs.</li> </ul>

opencc-by-sa-4.0Sep 2016View details →
zenodo44/100

Antisemitism on Twitter: A Dataset for Machine Learning and Text Analytics

<h1><strong><span><span>Dataset from the Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University:&nbsp; </span></span></strong></h1> <p>&nbsp;</p> <div> <div> <p><span><span>The </span><span>Social Media</span><span> &amp; Hate research lab at the Institute for the Study of Contemporary Antisemitism compiled this dataset using an annotation portal (Jikeli, Soemer, and Karali 2024), which was used to label tweets as either antisemitic or non-antisemitic, among other labels. Note that annotation was done on live data, including images and context, such as threads. All data was annotated by two experts, and all discrepancies were discussed</span><span> (Jikeli et al. 2023)</span><span>.</span></span><span>&nbsp;</span></p> </div> </div> <p><br><strong>Content: </strong></p> <p><span><span>This dataset&nbsp;</span><span>contains</span> <span>1</span><span>1</span><span>311</span><span> tweets </span><span>covering</span><span>&nbsp;a wide range of topics common in conversations about Jews, Israel, and antisemitism between January 2019 and </span><span>April 2023</span><span>.&nbsp;</span><span>The dataset consists of random samples of relevant keywords during this </span><span>time period</span><span>.</span><span>&nbsp;1,</span><span>953</span><span> tweets (1</span><span>7</span><span>%) </span><span>are antisemitic </span><span>according to </span><span>the IHRA definition of antisemitism.</span><span>&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> <div> <p><span><span>The distribution of tweets by year is as follows:</span><span>&nbsp;1499 (</span><span>13</span><span>%) from 2019, 371</span><span>2</span><span> (</span><span>33</span><span>%) from 2020, </span><span>2591</span><span> (2</span><span>3</span><span>%) from 2021</span><span>, 2644 from 2022 </span><span>(23%)</span> <span>and 865 </span><span>(8%)</span> <span>f</span><span>rom 2023</span><span>. </span><span>6365</span><span> (</span><span>56</span><span>%) </span><span>contain</span><span> the keyword "Jews,"</span><span> 4134 </span><span>(</span><span>3</span><span>7</span><span>%) include "Israel," 529 (</span><span>5</span><span>%) feature the derogatory term "</span><span>ZioNazi</span><span>*," and 283 (</span><span>3</span><span>%) use the slur "K---s." Some tweets may </span><span>contain</span><span> multiple keywords.&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>725</span><span> out of the </span><span>6365</span><span> tweets with the keyword "Jews" (11%) and </span><span>664</span><span> out of the </span><span>4134</span><span> tweets with the keyword "Israel" (1</span><span>6</span><span>%) were classified as antisemitic. 97 out of the 283 tweets using the antisemitic slur "K---s" (34%) are antisemitic.</span> <span>Interestingly, many tweets featuring the slur "K---s" actually </span><span>call out</span><span> its u</span><span>s</span><span>e.</span><span> In contrast, </span><span>the majority of</span><span> tweets </span><span>using</span><span> the derogatory term "</span><span>ZioNazi</span><span>*" are antisemitic, with 467 out of 529 (88%) being classified as such.&nbsp;</span></span><span>&nbsp;</span></p> </div> <p>&nbsp;</p> <p><strong>File Description:&nbsp;</strong></p> <div> <div> <p><span><span>The dataset is provided in a csv file format, with each row </span><span>representing</span><span> a single message, including replies, quotes, and retweets. The file </span><span>contains</span><span> the following columns:&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&nbsp;</span></span><span><span>&nbsp;</span><br></span><span><span>&lsquo;ID&rsquo;:</span> <span>Represents</span><span> the tweet ID.&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Username&rsquo;: </span><span>Represents</span><span> the username </span><span>that posted </span><span>the tweet</span><span>.&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Text&rsquo;: </span><span>Represents</span><span> the full text of the tweet (not pre-processed).</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;</span><span>CreateDate</span><span>&rsquo;: </span><span>Represents</span><span> the date </span><span>on which </span><span>the tweet was created</span><span>.&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Biased&rsquo;: </span><span>Represents</span><span> the label given by our annotations as to whether the tweet </span><span>is antisemitic or no</span><span>t</span><span>.</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Keyword&rsquo;: </span><span>Represents</span><span> the keyword that was used in the query. The keyword can be in the text, including </span><span>hashtags, </span><span>mentioned </span><span>users</span><span>, or the username</span><span> itself.</span><span>&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> </div> </div> <p>&nbsp;</p> <p>Licences&nbsp;</p> <p>Data is published under the terms of the "Creative Commons Attribution 4.0 International" licence (https://creativecommons.org/licenses/by/4.0)&nbsp;</p> <p>&nbsp;</p> <p>Acknowledgements&nbsp;</p> <p>We are grateful for the support of Indiana University&rsquo;s Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media &amp; Hate Research Lab at Indiana University&rsquo;s Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenh&auml;usler, Sophie von M&aacute;ri&aacute;ssy, Mabel Poindexter, Jenna Solomon, Clara Schilling, and Victor Tschiskale.&nbsp;</p> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Twitter hashtags time series used in the paper "Universality, criticality and complexity of information propagation in social media"

<pre>These files contain the time series and the associated hashtags we obtained by sampling Twitter for our paper &quot;Universality, criticality and complexity of information propagation on social media&quot;. The analysis is reported in <a href="https://arxiv.org/abs/2109.00116">https://www.nature.com/articles/s41467-022-28964-8</a> Please acknowledge the use of these data by citing the paper above. ################################# ################################# DATA ORGANIZATION We created a single zip file with all the time series and a single zip file with all the hashtags. There is a one-to-one correspondence between lines in the two files. ################################# ################################# FILES CONTENT As stated, here is a one-to-one correspondence between lines in the time series file and lines in the hashtags file, i.e., the hashtag stored in line X is the hashtag of the time series stored in line X. Time series are stored as follows: Ka t1 t2 t3 \n Kb t1 t2 t3 t4 t5 \n . . . Kn t1 t2 \n where: Ka, Kb,..., Kn is an integer specifying the number of events that compose the time series a, b,..., n respectively. In the example above we would have Ka=3, Kb=5, Kn=2. t1 t2 ... is the time series, i.e., a sequence of chronologically ordered interevent times. The last interevent time, in our implementation, represents the distance between the end of the temporal window and the last event time. It thus does not represent an event. As stated in the Supplemental Material of our paper, the temporal window ranges from 2019, October 1st to 2019, November 30th. </pre>

opencc-by-4.0Dec 2021View details →
zenodo44/100

TWikiL - Twitter Wikipedia Link Dataset

<p>The Twitter Wikipedia Link (TWikiL) dataset contains all Tweets posted on Twitter that contain a Wikipedia URL. The data was collected via Twitters academic research access and spans 15 years of Tweets from March&nbsp;2006 to January 2021. TWikiL comes in two versions: <strong>TWikiL_raw</strong> is a list of Tweet IDs in CSV format. <strong>TWikiL_curated</strong> is an SQLite database, which is a curated version of TWikiL containing only links to Wikipedia articles. The curated version has been augmented with the language edition that the URL in the Tweet links to, the Wikidata identifier and a Wikipedia topic category.&nbsp;<br> <br> TWikiL raw contains&nbsp;44,945,098 Tweet IDs<br> TWikiL curated contains&nbsp;35,252,782 URLs/Wikidata concepts with&nbsp;34,543,612 unique Tweets and&nbsp;474,577 Tweets linking to multiple Wikipedia articles.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

A Multilingual Dataset of COVID-19 Vaccination Attitudes on Twitter

<p>This dataset consists of the IDs of 2,198,090 tweets collected from Western Europe,&nbsp;of which 17,934 are annotated with&nbsp;labels indicating the originators&#39; affective vaccination stances,&nbsp;including Positive (PO), Negative (NG), Positive but dissatisfaction (PD), Neutral (NE) and Off-topic (OT).</p> <p>all_tweets.txt contains all the ids of the collected tweets, annotated_tweets.txt&nbsp;contains the ids of the annotated tweets and the categories they are annotated to.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Twitter Account Dataset - 10,000 accounts tested through the Botometer API

<p>This dataset contains 10,420 accounts extracted from Twitter from tweets containing the word &quot;vaccine&quot;. We verified each one through the Botometer API.</p> <p>Includes:</p> <ul> <li>Account ID</li> <li>Echo Chamber Score (0-5)</li> <li>Fake Follower Score (0-5)</li> <li>Financial Score (0-5)</li> <li>Self Declared Score (0-5)</li> <li>Spammer Score (0-5)</li> <li>Other Score (0-5)</li> <li>Overall Score (0-5)</li> </ul> <p>Those appearing with an &#39;X&#39; correspond to deleted, modified or private accounts.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Public Dataset for "Did State-sponsored Trolls Shape the 2016 US Presidential Election Discourse? Quantifying Influence on Twitter"

<p>Dataset for the &quot;Did State-sponsored Trolls Shape the 2016 US Presidential Election Discourse? Quantifying Influence on Twitter&quot; paper.&nbsp;</p> <p>The full text of the paper can be found&nbsp;<a href="https://zenodo.org/record/4699959#.YngKatNBy3K">here</a>.</p> <p>The folder &quot;Tweet_IDs&quot; contains the complete list of the 152,514,929 tweet IDs (together with their timestamps) which we used for the analysis in the study: &quot;Did State-sponsored Trolls Shape the 2016 US Presidential Election Discourse? Quantifying Influence on Twitter&quot;</p> <p>by Nikos Salamanos, Michael J. Jensen, Costas Iordanou and Michael Sirivianos</p> <p>We have split the tweets into separate .zip files based on the date listed in their timestamps.</p> <p>The crawling took place from September 21 to November 7, 2016 (47 days; we did not collect data on 02/10/2016).</p> <p>Each &quot;tweet_day_X.zip&quot; file contains the file &quot;tweet_day_X.csv&quot;, where X in [1,2,...,47]. For instance, the file &quot;tweets_day_1.zip&quot; contains the tweets of the 1st day: 09/21/2016.</p> <p>Please cite the paper in any published work that uses any of these resources.&nbsp;</p> <p>@misc{nikos_salamanos_2021_4699959,<br> &nbsp; author &nbsp; &nbsp; &nbsp; = {Nikos Salamanos and<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Michael J. Jensen and<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Costas Iordanou and<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Michael Sirivianos},<br> &nbsp; title &nbsp; &nbsp; &nbsp; &nbsp;= {{Did State-sponsored Trolls Shape the 2016 US&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Presidential Election Discourse? Quantifying<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Influence on Twitter}},<br> &nbsp; month &nbsp; &nbsp; &nbsp; &nbsp;= apr,<br> &nbsp; year &nbsp; &nbsp; &nbsp; &nbsp; = 2021,<br> &nbsp; publisher &nbsp; &nbsp;= {Zenodo},<br> &nbsp; version &nbsp; &nbsp; &nbsp;= 3,<br> &nbsp; doi &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;= {10.5281/zenodo.4699959},<br> &nbsp; url &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;= {https://doi.org/10.5281/zenodo.4699959}<br> }</p>

opencc-by-4.0May 2022View details →
zenodo44/100

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

<p>We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria&mdash;Hausa, Igbo, Nigerian-Pidgin, and Yor&ugrave;b&aacute;&mdash;consisting of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), including a significant fraction of code-mixed tweets.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Dataset: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter

<p>This is the dataset, trained model, and software companion for the paper titled: Characterizing Anti-Asian Rhetoric During The COVID-19 Pandemic: A Sentiment Analysis Case Study on Twitter accepted for the Workshop on Data for the Wellbeing of Most Vulnerable of the&nbsp;ICWSM 2022 conference.</p> <p>The COVID-19 pandemic has shown a measurable increase in the usage of sinophobic comments or terms on online social media platforms. In the United States, Asian Americans have been primarily targeted by violence and hate speech stemming from negative sentiments about the origins of the novel SARS-CoV-2 virus. While most published research focuses on extracting these sentiments from social media data, it does not connect the specific news events during the pandemic with changes in negative sentiment on social media platforms. In this work we combine and enhance publicly available resources with our own manually annotated set of tweets to create machine learning classification models to characterize the sinophobic behavior. We then applied our classifier to a pre-filtered longitudinal dataset spanning two years of pandemic related tweets and overlay our findings with relevant news events.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Russo-Ukrainian War: Prediction and explanation of Twitter suspension

<p>Dataset used in research paper : &quot;Russo-Ukrainian War: Prediction and explanation of Twitter<br> suspension&quot;. Contain extracted features for Twitter users, correlated to discussion of Russo-Ukrainian war. Contain multiple feature categories. Source code of dataset usage in Machine Learning approach is available at GitHub: https://github.com/alexdrk14/TwitterSuspension .</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

SSIX BREXIT Twitter Annotated Data Set

<p><strong>SSIX BREXIT Gold Standard</strong></p> <p>This repository contains the BREXIT Twitter Gold Standard produced by the&nbsp;SSIX Project <a href="https://ssix-project.eu/">https://ssix-project.eu/</a>.</p> <p>Only a sample is available here,&nbsp;to rebuild the full dataset,&nbsp;follow the instructions on the SSIX Project code&nbsp;repository:&nbsp;</p> <p><a href="https://bitbucket.org/ssix-project/brexit-gold-standard">https://bitbucket.org/ssix-project/brexit-gold-standard</a></p>

opencc-by-sa-4.0Dec 2016View details →
zenodo44/100

MMoveT15: A Twitter Dataset for Extracting and Analysing Migration-Movement Data of the European Migration Crisis 2015

<p>In the 2015 migration crisis thousands of refugees and migrants crossed the border to Hungary, Austria and Germany. The movements of these people are reflected in social media, especially on Twitter. We present a dataset of 3275 Tweets form the months September and October 2015. These Tweets are annotated regarding their relevance to the quantitative movement of refugees/migrants into Hungary, Austria and Germany. We present this dataset for a posterior analysis of the 2015 migration crisis or as a basis for an early warning or forecasting system</p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

A Twitter Dataset for Spatial Infectious Disease Surveillance

<p>Dengue is a mosquito-borne viral disease which infects millions of people every year, specially in developing countries. Some of the main challenges facing the disease are reporting risk indicators and rapidly detecting outbreaks. Traditional surveillance systems rely on passive reporting from health-care facilities, often ignoring human mobility and locating each individual by their home address. Yet, geolocated data are becoming commonplace in social media, which is widely used as means to discuss a large variety of health topics, including the users&#39; health status. In this dataset paper, we make available two large collections of dengue related labeled Twitter data. One is a set of tweets available through the Streaming API using the keywords dengue and aedes from 2010 to 2016. The other is the set of all geolocated tweets in Brazil during the year of 2015 (available also through the Streaming API). We detail the process of collecting and labeling each tweet containing keywords related to dengue in one of 5 categories: personal experience, information, opinion, campaign, and joke. This dataset can be useful for the development of models for spatial disease surveillance, but also scenarios such as understanding health-related content in a language other than English, and studying human mobility.</p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

Forschungsdaten/Visualisierungen zu: "Microblogging in den Informationswissenschaften - Quantitative Untersuchungen exemplarischer Communities auf Twitter"

<p>Dieses Datenset enth&auml;lt zus&auml;tzliche Forschungsdaten zur Bachelorarbeit <a href="https://opus4.kobv.de/opus4-fhpotsdam/frontdoor/index/index/docId/2340">&quot;Microblogging in den Informationswissenschaften - Quantitative Untersuchungen exemplarischer Communities auf Twitter&quot;</a>. Eine genaue Erl&auml;uterung der einzelnen Dateien findet in der Arbeit selbst statt. Die hier enthaltenen Personendaten wurden nicht anonymisiert, enthalten jedoch rein &ouml;ffentlich zug&auml;ngliche Informationen.</p> <p>Zus&auml;tzlich sind hier einzelne Visualisierungen aus der Arbeit im PDF-Format enthalten.</p>

opencc-by-4.0May 2019View details →
zenodo44/100

The complete corpus of #COVID-19 Twitter dataset

<p><br> COVID-19 pandemic initiated over a year ago continues to spread around the globe and the ongoing research regarding COVID-19 is on a continues growth as well. The online discourse on social media regarding COVID-19 has been growing along with the timeline of the pandemic.</p> <p>Open data on Twitter have been released and offer the research community the opportunity for new findings and resolving this new threat.&nbsp;&nbsp;In this dataset, we open a corpus of Twitter&#39;s data from March 2020 till today, that is being updated every day based on the two most important hashtags regarding COVID-19. &nbsp;This dataset will offer the research community the opportunity to explore the social extensions of this pandemic including topic analysis, hate speech sentiment analysis, regarding either the opinion of the users on the pandemic, the comments on the public discourse, or the vaccination releases.&nbsp; The dataset has been collected by retrieving all the tweets that contain the hashtags: #coronavirus and #COVID19 &nbsp;including approximately 208M tweets for hashtags #coronavirus and 392M tweets for hashtag #COVID-19, resulting in a total of 600M tweets.&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Dataset of traffic accidents reported on Twitter Bogotá Colombia

<p><strong>1 Classification Dataset</strong></p> <p>This dataset for the classification model contains 3,804 tweets, where 1,902 are related to traffic accident reports (TA, positive class) and 1,902 are unrelated (NTA, negative class).</p> <p>For training the tweet classification model, a collaborative labeling strategy was designed. Here, 30 people labeled data according to the instructions given. Each participant had to evaluate a tweet to manually classify it into one of three categories defined as: traffic accident related, unrelated and don&acute;t know/no response. Each tweet was evaluated by 3 participants. The correct label was selected by voting; the 3 people must agree on the selected label, otherwise the tweet was excluded from training. This process took a month and required the development and deployment of a web application.</p> <p><strong>2 NER Dataset (Named Entity Recognition)</strong></p> <p>For the entity recognition model training, a sample of the filtered tweets resulting from the previous classification phase was taken. 1,340 tweets were extracted, where 800 are from &ldquo;unofficial&rdquo; users, almost 60% of the sample. These tweets were user reports on traffic incident occurred in Bogota from October 2018 to July 2019, including other tweets that contained some location references such as reports on the state of road infrastructure; some tweets from the years 2016 and 2017 were also included. Although these posts were not related to accidents per se, they were selected because they contained location information. The purpose was to train a model that would recognize these entities, because a classifier of accident-related tweets was previously created. Additionally, the dataset was split, reserving 1,072 tweets for training and 268 for evaluation.</p> <p>This dataset was manually labeled using the IOB (Inside-outside- beginning) format. The labeling tool called Brat Annotation Tools was used for this task. The labels defined are Location, which refers to the location of the report; and Time, which refers to the time or date of the incident. Accordingly, 5 labels were generated: B-loc, I-loc, B-time, I-time and O. The O label refers to Others.</p> <p><strong>3 Traffic accident Twitter geolocation</strong></p> <p>A dataset with 26362 traffic accident tweets with the coordinates of the incident and the date of publication.</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record