Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
307
datasets available to search
ShareScore release 0.7.1
Dataset results
307 results for “twitter”
Russo-Ukrainian War: Prediction and explanation of Twitter suspension
<p>The dataset utilized in the research paper: "Russo-Ukrainian War: Prediction and explanation of Twitter suspension" accepted to ASONAM 2023 conference. The provided dataset contains multiple extracted feature categories based on the Twitter dataset collected during the Russo Ukrainian War. The dataset dose not contain any private user information since user and tweet IDs are removed.</p>
Polarized and nonpolarized Twitter networks from the 2019 Finnish Parliamentary Elections
<p><strong>Polarized and nonpolarized Twitter networks from the 2019 Finnish Parliamentary Elections</strong></p> <p>This dataset includes 183 Twitter retweet networks collected during the 2019 Finnish Parliamentary Elections.</p> <p>The first 150 networks are built around single hashtags, such as #police, #nature, and #immigration. The remaining 33 networks are constructed using a combination of hashtags focused on specific topics like climate change and economic policy.</p> <p>Each filename consists of two parts: the first part indicates whether the network is based on a single hashtag (in lowercase) or a set of hashtags (in uppercase). The second part represents the tweet period.</p> <ul> <li> <p>"p1" corresponds to the pre-election period (March 1 to April 14).</p> </li> <li> <p>"p2" corresponds to the inter-election period (April 15 to May 26).</p> </li> <li> <p>"p3" corresponds to the post-election period (May 27 to July 31).</p> </li> </ul> <p>The nodes in the networks represent anonymized Twitter accounts, and directed ties indicate retweet endorsements on specific topics. Each file contains three columns: retweeter, retweeted, and weight.</p> <p>Please see the references for more details.</p> <p>Network labels, whether they are labeled as controversial, and whether they are based on single or multiple hashtags, can be found in the "networks_info.csv" file.</p> <p>Importantly, the dataset does not contain any identifying information or original raw data from the Twitter platform. Anonymization was achieved by shuffling the order of unique nodes across all networks and assigning each node a new identifier (ID). These new IDs were then applied to the edgelists to obtain the anonymized version.</p> <p>Kindly ensure to reference the original article(s) when utilizing this dataset.</p> <p>Chen, T. H. Y., Salloum, A., Gronow, A., Ylä-Anttila, T., & Kivelä, M. (2021). Polarization of climate politics results from partisan sorting: Evidence from Finnish Twittersphere. <em>Global Environmental Change</em>, <em>71</em>, 102348. <a href="https://doi.org/10.1016/j.gloenvcha.2021.102348">https://doi.org/10.1016/j.gloenvcha.2021.102348</a></p> <p>Salloum, A., Chen, T. H. Y., & Kivelä, M. (2022). Separating polarization from noise: comparison and normalization of structural polarization measures. <em>Proceedings of the ACM on human-computer interaction</em>, <em>6</em>(CSCW1), 1-33. <a href="https://doi.org/10.1145/3512962">https://doi.org/10.1145/3512962</a></p>
Twitter Crawling for "Viral or Heboh" News
<p>The data we take is about tweet form 10 official account of news portal on twitter that mentioned word "Viral" or "Heboh". The dataset contain 5 columns and more than 500 tweet sorted by time.</p>
Twitter Dataset for "Will You Take the Knee? Italian Twitter Echo Chambers' Genesis During EURO 2020"
<p>Echo chambers can be described as situations in which individuals encounter and interact only with viewpoints that confirm their own, thus moving, as a group, to more polarized and extreme positions. Recent literature mainly focuses on characterizing such entities via static observations, thus disregarding their temporal dimension. In this work, distancing from such a trend, we study, at multiple topological levels, echo chambers genesis related to the social discussions that took place in Italy during the EURO 2020 Championship. Our analysis focuses on a well-defined topic (i.e., BLM/racism) discussed on Twitter during a perfect temporally bound (sporting) event. Such characteristics allow us to track the rise and evolution of echo chambers in time, thus relating their existence to specific episodes.</p>
Twitter dataset of flood-related images for September 2021, Thailand and June/July 2021, Nepal floods
<p>Twitter dataset related to flood events onsets in Thailand and Nepal, focused on September 26/27, 2022, June 16/17 2021 and July 01/02 2021. The dataset has been processed with a VisualCit pipeline in order to automatically filter a relevant subset of posts through automated image analysis, using deep learning techniques. The posts were then geolocated using the CIME algorithm. Additional information about the data collection and data processing are described in <a href="http://arxiv.org/abs/2202.12014">http://arxiv.org/abs/2202.12014</a></p>
Twitter Dataset for ≈15,000 accounts over a 1 day period
<p>This is almost 15,000 accounts on Twitter. Accounts were collected from a 1% sampled stream over a 24-hour period, from 2022-07-02T04:18:23.000Z to 2022-07-03T04:18:23.000Z.</p> <p>This data includes data from labeled data from https://doi.org/10.5281/zenodo.2653137, with an additional random sample of accounts from the stream to get to 15,000 accounts. Private profiles were removed.</p> <p>This data was collected with the intention of using it for unsupervised machine learning. You are free to do what you want with proper citation.</p>
Dia-Pol: A large scale BlackLivesMatter and MeToo Twitter dataset
<p>This dataset (tweets_id_list.json) contains 258609 number of tweets sent in English extracted from Twitter API using the query word “#blacklivesmatter.” The dataset spans the period from 2020-01-01 to 2021-12-31 and was retrieved on 2022-06-10.</p>
SenTopX: A Benchmark Twitter Dataset for User Sentiment on Various Topics
<p>This is a longitudinal Twitter dataset of 143K users during the period 2017-2021. The following is the detail of all the files:</p> <ul> <li><a href="11243662" target="_blank" rel="noopener noreferrer">SenTopX_userIDs.txt</a>: contains user IDs of 143K Twitter users.</li> <li><a href="../api/records/11243662/draft/files/userIDs_tweetIDs.zip/content" target="_blank" rel="noopener noreferrer">userIDs_tweetIDs.zip</a>: contains Tweet IDs of users, the name of the file is the user ID and the file contains the list of all the tweet IDs.</li> <li><a href="../api/records/11243662/draft/files/users_16_perspective_toxicity_scores.csv/content" target="_blank" rel="noopener noreferrer">users_16_perspective_toxicity_scores.csv</a> contains user IDs and 16 median Perspective API scores, the vector is shared as mean, median, and Gini Index of scores calculated over all tweets of a user.</li> <li><a href="../api/records/11243662/draft/files/LDAvis_top30_words_for_extracted_topics.csv/content" target="_blank" rel="noopener noreferrer">LDAvis_top30_words_for_extracted_topics.csv</a> contains the top 30 most relevant words extracted from each topic extracted by tweet-level topic modeling using the BERTweet topic model.</li> <li><a href="../api/records/11243662/draft/files/topic_modelling_statistics_per_user.csv/content" target="_blank" rel="noopener noreferrer">topic_modelling_statistics_per_user.csv</a> contains important and relevant statistics related to topic modeling results: <ul> <li> <p>1. user: This column represents the identifier for the user. Each row in the CSV corresponds to a specific user, and this column helps to track and differentiate between the users.</p> <p>2. avg_topic_probability: This column contains the average probability of the topics for each user calculated across all of the tweets in order to compare users in a meaningful way. It represents the average likelihood that a particular user discusses various topics over the observed period.</p> <p>3. maximum_topic_avg: This column holds the value of the highest average probability among all topics for each user. It indicates the topic that the user most frequently discusses, on average.</p> <p>4. index_max_avg_topic_probability_200: This column specifies the index or identifier of the topic with the highest average probability out of 200 possible topics. It shows which topic (out of 200) the user discusses the most.</p> <p>5. global_avg: This column includes the global average probability of topics across all users. It provides a baseline or overall average topic probability that can be used for comparative purposes.</p> <p>6. max_global_avg: This column contains the maximum global average probability across all topics for all users. It identifies the most discussed topic across the entire user base.</p> <p>7. index_max_global_avg: This column shows the index or identifier of the topic with the highest global average probability. It indicates which topic (out of 200) is the most popular across all users.</p> <p>8. entropy_200_topic: This column represents the entropy of the topics for each user, calculated over 200 topics. Entropy measures the diversity or unpredictability in the user's discussion of topics, with higher entropy indicating more varied topic discussion.</p> <p>In summary, these columns are used to analyze the topic engagement and preferences of users on a platform, highlighting the most frequently discussed topics, the variability in topic discussions, and how individual user behavior compares to overall trends.</p> </li> </ul> </li> </ul>
#Élysée2017fr: The 2017 French Presidential Campaign on Twitter
<p># README</p> <p>This archive contains the #Élysée2017fr dataset.</p> <p>(Initially published at https://web.archive.org/web/20200530171644if_/https://dataverse.mpi-sws.org/dataverse/icwsm18 on June 24, 2018. This dataverse being defunct now, we repost on Zenodo)</p> <p><br> ## Content</p> <p>### keywords.csv<br> The keywords used to collect the initial dataset, each presented with the start and stop dates of use (date format: YYYY-MM-DD).</p> <p><br> ### profiles_annotations.csv<br> The manual profiles annotations. The file contains the following columns:</p> <p>#### FROM_USER_ID<br> The profile's id used by Twitter</p> <p>#### PROFILE_NATURE<br> "**individual**" if the profile is managed by a single person, else "**non individual**".<br> The "**non individual**" label is itself divided in 3 subcategories:<br> - "**political**" for profiles of political parties or associations, and profiles representing groupes of militants.<br> - "**media**" for profiles of media outlets.<br> - "**other**" for profiles not included in the previous categories.</p> <p>#### PARTY<br> The profile's political affiliation(s), indicated as the shortcut for the political party:<br> - "**fi**": France Insoumise (far-left)<br> - "**ps**": Parti Socialiste (left)<br> - "**em**": En Marche ! (center)<br> - "**lr**": Les Républicains (right)<br> - "**fn**": Front National (far-right)<br> - **null**: no political affiliation</p> <p>When a profile has 2 affiliations, they are separated by a slash (ex: "ps/fi"). </p> <p>#### MEDIA_PROFESSIONAL<br> *For individual profiles only.*<br> Indicates if the profile's owner self-identify as a media professional (journalist, editorialist, ...)</p> <p>#### SEX<br> *For individual profiles only.*<br> Indicates the sex of the profile's owner:<br> - "**m**": male<br> - "**f**": female<br> - **null**: undetermined or other</p> <p><br> ### posts_ids_*<br> Files containing the tweets and retweets ids, divided according to the political affiliation of their authors for more flexibility.<br> - **posts_ids_fi.csv**: Tweet ids for profiles affiliated to France Insoumise (far-left)<br> - **posts_ids_ps.csv**: Tweet ids for profiles affiliated to Parti Socialiste (left)<br> - **posts_ids_em.csv**: Tweet ids for profiles affiliated to En Marche ! (center)<br> - **posts_ids_lr.csv**: Tweet ids for profiles affiliated to Les Républicains (right)<br> - **posts_ids_fn.csv**: Tweet ids for profiles affiliated to Front National (far-right)<br> - **posts_ids_multi_affiliations.csv**: Tweet ids for profiles affiliated to more than one party<br> - **posts_ids_indetermined.csv**: Tweet ids for profiles not affiliated to any party</p> <p>Each file contains one tweet id per line.</p> <p><br> ### networks_*<br> Files containing the mention and retweet networks, in NCOL and GraphML format.</p> <p>The NCOL files contains the directed weighted edges between profiles, one per line, in the following format:<br> profile1_twitter_id profile2_twitter_id edge_weight</p> <p>The GraphML files contains the directed weighted edges between profiles, as well as all the profiles annotations presented in *profiles_annotations.csv*. They can be opened using a graph visualisation software like Gephi.</p> <p> </p> <p>## How to get tweets from ids<br> You can use various tools to help you get tweets from their ids, we suggest the following:<br> - DMI-TCAT: https://github.com/digitalmethodsinitiative/dmi-tcat<br> - Twarc: https://github.com/DocNow/twarc</p> <p> </p> <p>## How to cite this work<br> Fraisier Ophélie, Cabanac Guillaume, Pitarch Yoann, Besançon Romaric, Boughanem Mohand. 2018. #Élysée2017fr: the French Presidential Election on Twitter. In International Conference on Weblogs and Social Media. https://aaai.org/ocs/index.php/ICWSM/ICWSM18/paper/view/17821 (https://hal.archives-ouvertes.fr/hal-02319715)</p>
[Dataset] Analysis of ego-networks of two CS-related Twitter accounts
<p><strong>Explanation/Overview:</strong></p> <p>This is the dataset for the analyses and results on Twitter Ego-Networks of two CS-related accounts (<a href="https://twitter.com/EuCitSci">@EuCitSci</a> & <a href="https://twitter.com/SciStarter">@SciStarter</a>). The username have been anonymised.</p> <p><strong>Purpose:</strong></p> <p>The purpose of this dataset is to provide the basis to reproduce the results reported in the associated deliverable. As such, it <strong>does not</strong> represent <strong>raw data</strong>, but rather files that already include certain analysis steps (like calculated degrees or other SNA-related measures), ready for analysis, visualisation and interpretation with R or any other network visualisation software (e.g., Gephi). The edges represent the follow relation and were retrieved using the Twitter API. All usernames except those of the two ego-accounts were anonymised by assigning a random number to each node. Due to the size of the network, we do not include any <code>.gexf </code>or <code>.gml </code>files in this upload, but rather resort to node and edge lists in the <code>.json</code> format</p> <p><strong>Relatedness:</strong></p> <p>The networks are the ego-networks for two related public accounts that are associated with citizen science (<a href="https://twitter.com/EuCitSci">@EuCitSci</a> & <a href="https://twitter.com/SciStarter">@SciStarter</a>).</p> <p><strong>Content:</strong></p> <p>In this Zenodo entry, two files can be found.</p> <ul> <li><code>edges.json</code></li> </ul> <p>Represents the edge list, with the columns:</p> <table> <tbody> <tr> <td><code>source</code></td> <td><code>target</code></td> <td><code>wherefrom</code></td> <td><code>account</code></td> </tr> <tr> <td>EuCitSci</td> <td>23929</td> <td>friendslist</td> <td>EuCitSci</td> </tr> </tbody> </table> <p><code>Source</code> and <code>target</code> are the necessary columns for the network creation and in this example indicate that EuCitSci follows 23929. <code>wherefrom </code>indicates the origin of this relation in the crawling (i.e., whether it was retrieved using the friends or follower list) and <code>account </code>indicates the ego-account it belongs to. Thus, the edges can also be separated using this <code>account </code>attribute as they represent two distinct networks.</p> <ul> <li><code>nodes.json</code></li> </ul> <p>Represents the nodes in the networks. The following data fields are contained:</p> <table> <tbody> <tr> <td><code>username</code></td> <td><code>CS</code></td> <td><code>followers_count</code></td> <td><code>friends_count</code></td> <td>...</td> </tr> <tr> <td>1116</td> <td>CS</td> <td>689</td> <td>514</td> <td>...</td> </tr> </tbody> </table> <table> <tbody> <tr> <td>...</td> <td><code>favourites_count</code></td> <td><code>listed_count</code></td> <td><code>statuses_count</code></td> <td>...</td> </tr> <tr> <td>...</td> <td>2601</td> <td>18</td> <td>2141</td> <td>...</td> </tr> </tbody> </table> <table> <tbody> <tr> <td>...</td> <td><code>degree</code></td> <td><code>in_degree</code></td> <td><code>out_degree</code></td> <td>...</td> </tr> <tr> <td>...</td> <td>21</td> <td>5</td> <td>16</td> <td>...</td> </tr> </tbody> </table> <table> <tbody> <tr> <td>...</td> <td><code>reciprocity</code></td> <td><code>account</code></td> </tr> <tr> <td>...</td> <td>0.48</td> <td>EuCitSci</td> </tr> </tbody> </table> <p> </p> <p><code>Username </code>represents the numerical and anonymised username, <code>CS</code> the community-membership. The different counts (e.g., <code>followers_count</code>) indicate the number of followers the user had at the time of the retrieval by the Twitter API. <code>degree</code> refers to the degree in the network (similarly the <code>in-</code> and <code>out_degree</code>), while <code>reciprocity</code> refers to the number of mutual edges in respect to the total number of edges per node (see <a href="https://networkx.org/documentation/stable/reference/algorithms/generated/networkx.algorithms.reciprocity.reciprocity.html">here</a>). <code>Account</code> is similar as described above.</p> <p> </p> <p><strong>Grouping:</strong></p> <p>The data is grouped according the ego-account it is associated to, as can be read above (i.e., the <code>account</code> attribute).</p>
VaxxHesitancy: A Dataset for Studying Hesitancy Towards COVID-19 Vaccination on Twitter
<p>We create a publicly available dataset of over 3,100 COVID-19 vaccine-related tweets labeled as one of four stance categories: <em>pro-vaxx, anti-vaxx</em>, <em>vaxx-hesitant</em>,<em> or irrelevant</em>.</p> <p><strong>***</strong></p> <p><strong>Please use the V2 version.</strong></p> <p><strong>***</strong></p> <p>We split our dataset into two separate files:</p> <p>(1) VaccineHesitancy_train_v2.csv (Single + Double annotated)</p> <p>(2) VaccineHesitancy_test.csv (Double annotated)</p> <p>We present the details of this dataset here:</p> <p>VaxxHesitancy: A Dataset for Studying Hesitancy Towards COVID-19 Vaccination on Twitter (ICWSM 2023)</p> <p><strong>Our Pre-trained model</strong> (GateNLP/covid-vaccine-twitter-bert) : https://huggingface.co/GateNLP/covid-vaccine-twitter-bert</p> <p><strong>Paper</strong>: https://ojs.aaai.org/index.php/ICWSM/article/view/22213/21992</p> <p> </p> <pre>@inproceedings{mu2023vaxxhesitancy, title={VaxxHesitancy: A Dataset for Studying Hesitancy Towards COVID-19 Vaccination on Twitter}, author={Mu, Yida and Jin, Mali and Grimshaw, Charlie and Scarton, Carolina and Bontcheva, Kalina and Song, Xingyi}, booktitle={Proceedings of the International AAAI Conference on Web and Social Media}, volume={17}, pages={1052--1062}, year={2023} } </pre> <p> </p> <p> </p> <p> </p>
Undirected Node Attributed Social Network Graph of Twitter Users interested in plastic pollution - created in the framework of the PlasticTwist project
<p>This dataset has been created in the framework of the Plastic Twist project (<a href="https://ptwist.eu/">Ptwist</a>) and more specifically using the Ptwist crowdsourcing application (<a href="https://crowdsourcing.plastictwist.com/">crowdsourcing.plastictwist.com/</a>). We are sharing the edge list and specific node attributes (hashtags) of Twitter users posting about plastic pollution. The dataset can be used for community detection,clustering, node importance, influence maximization tasks, etc. Each user is represented by a unique integer which has nothing to do with the official Twitter user ID. The dataset contains three (3) files: </p> <ul> <li>ptwist.edgelist: A list containing all the 1,362,863 edges between the users. When loaded they create an undirected graph of 800K+ users.</li> <li>node_attributes.txt: This file contains information about the hashtags used by each user. (e.g. "652003": ["SingleUsePlastic"] -> user 6529003 has used the hashtag SingleUsePlastic) </li> <li>annotated_graph: A pickle file which, when loaded, returns a <a href="https://networkx.github.io/">NetworkX</a> node attributed undirected graph.</li> </ul> <p> </p> <p> </p>
Public Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior"
<p>Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found <a href="https://arxiv.org/pdf/1802.00393.pdf">here</a>. </p> <p>The dataset provided here includes an updated version of the original dataset, with ~100k tweets annotated using the CrowdFlower platform: </p> <ul> <li> <p>hatespeech_id_label_PUBLIC_100K.csv: contains ~100K rows, where every row consists of a unique Tweet ID. </p> </li> <li> <p>hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K rows, where every row consists of the tweet text, its label according to majority annotation and the number of majority annotators. Available only <a href="https://zenodo.org/record/3706866#.Xmkh6i97FQI">here</a>.</p> </li> <li> <p>retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Available only <a href="https://zenodo.org/record/3706866#.Xmkh6i97FQI">here</a>.</p> </li> </ul> <p> </p> <p>UPDATE: It has come to our understanding that a number of the tweets are not available anymore for download on Twitter. Therefore, we provide <a href="https://zenodo.org/record/3706866#.YYLG6S8RqjQ">here </a>the hatespeech_text_label_vote_RESTRICTED_100K file with the full ~100K tweet texts, their associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). </p> <p>Please cite the paper in any published work that uses any of these resources. </p> <p>@inproceedings{founta2018large, <br> title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior}, <br> author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas}, <br> booktitle={11th International Conference on Web and Social Media, ICWSM 2018}, <br> year={2018}, <br> organization={AAAI Press} <br> } </p> <p>For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy </p>
Digital Narratives of Covid-19: a Twitter Dataset
<p>We are releasing a Twitter dataset connected to our project <a href="https://covid.dh.miami.edu/"><em>Digital Narratives of Covid-19</em> </a>(DHCOVID) that -among other goals- aims to explore during one year (May 2020-2021) the narratives behind data about the coronavirus pandemic.</p> <p>In this first version, we deliver a Twitter dataset organized as follows:</p> <ul> <li>Each folder corresponds to daily data (one folder for each day): YEAR-MONTH-DAY</li> <li>In every folder there are 9 different plain text files named with "dhcovid", followed by date (YEAR-MONTH-DAY), language ("en" for English, and "es" for Spanish), and region abbreviation ("fl", "ar", "mx", "co", "pe", "ec", "es"): <ol> <li>dhcovid_YEAR-MONTH-DAY_es_fl.txt: Dataset containing tweets geolocalized in South Florida. The geo-localization is tracked by tweet coordinates, by place, or by user information.</li> <li>dhcovid_YEAR-MONTH-DAY_en_fl.txt: We are gathering only tweets in English that refer to the area of Miami and South Florida. The reason behind this choice is that there are multiple projects harvesting English data, and, our project is particularly interested in this area because of our home institution (University of Miami) and because we aim to study public conversations from a bilingual (EN/ES) point of view.</li> <li>dhcovid_YEAR-MONTH-DAY_es_ar.txt: Dataset containing tweets from Argentina.</li> <li>dhcovid_YEAR-MONTH-DAY_es_mx.txt: Dataset containing tweets from Mexico.</li> <li>dhcovid_YEAR-MONTH-DAY_es_co.txt: Dataset containing tweets from Colombia.</li> <li>dhcovid_YEAR-MONTH-DAY_es_pe.txt: Dataset containing tweets from Perú.</li> <li>dhcovid_YEAR-MONTH-DAY_es_ec.txt: Dataset containing tweets from Ecuador.</li> <li>dhcovid_YEAR-MONTH-DAY_es_es.txt: Dataset containing tweets from Spain.</li> <li>dhcovid_YEAR-MONTH-DAY_es.txt: This dataset contains all tweets in Spanish, regardless of its geolocation.</li> </ol> </li> </ul> <p>For English, we collect all tweets with the following keywords and hashtags: covid, coronavirus, pandemic, quarantine, stayathome, outbreak, lockdown, socialdistancing. For Spanish, we search for: covid, coronavirus, pandemia, quarentena, confinamiento, quedateencasa, desescalada, distanciamiento social.</p> <p>The corpus of tweets consists of a list of Tweet Ids; to obtain the original tweets, you can use "<a href="https://github.com/DocNow/hydrator">Twitter hydratator</a>" which takes the id and download for you all metadata in a csv file.</p> <p>We started collecting this Twitter dataset on April 24th, 2020 and we are adding daily data to our GitHub repository. There is a detected problem with file 2020-04-24/dhcovid_2020-04-24_es.txt, which we couldn't gather the data due to technical reasons.</p> <p>For more information about our project visit <a href="https://covid.dh.miami.edu/">https://covid.dh.miami.edu/</a></p> <p>For more updated datasets and detailed criteria, check our GitHub Repository: <a href="https://github.com/dh-miami/narratives_covid19/">https://github.com/dh-miami/narratives_covid19/</a></p>
#retweetthe8th: twitter dataset from the 2018 Referendum to repeal the 8th Amendment of the Constitution of Ireland
<p>This dataset contains the tweet ids of 2,108,782 tweets related to the referendum to repeal the 8th Amendment of the Constitution of Ireland and replace it with the Thirty-sixth Amendment of the Constitution of Ireland on May 25, 2018. </p> <p>They were collected between March 9th, 2018 and May 30th, 2018 from the Twitter filter stream API using Twarc (<a href="https://github.com/DocNow/twarc">https://github.com/DocNow/twarc</a>). The set of terms that were used for the search were: #repealthe8th, #together4yes, #8thref, #savethe8th, #hometovote, #togetherforyes, #repealedthe8th, #loveboth, #voteyes, #voteno, #lovebothvoteno, #repealtheeighth.</p> <p>Note that the terms changed during the course of data collection and that the searches were not exhaustive.</p> <p>The GET statuses/lookup method supports retrieving the complete tweet for a tweet id (known as hydrating). Twarc can also be used to hydrate tweets.</p> <p>Per Twitter’s Developer Policy, tweet ids may be publicly shared for academic purposes; tweets may not.</p> <p>There are several practical reasons to leave the retweets, as it allows researchers to trace important tweets and their dissemination. </p> <p>However, for researchers that might be interested in NLP tasks for which retweets are not required, a set of 411,213 tweet ids without their retweets is also included in the zip file.</p> <p>Questions about this dataset can be sent to Emmet Ó Briain: <a href="mailto:emmet@quiddity.ie">emmet@quiddity.ie</a> .</p> <p>(24-05-20)</p>
CMU-MisCov19: A Novel Twitter Dataset for Characterizing COVID-19 Misinformation
<p>From conspiracy theories to fake cures and fake treatments, COVID-19 has become a hot-bed for the spread of misinformation online. It is more important than ever to identify methods to debunk and correct false information online. Detection and characterization of misinformation requires an availability of annotated datasets. Most of the published COVID-19 Twitter datasets are generic, lack annotations or labels, employ automated annotations using transfer learning or semi-supervised methods, or are not specifically designed for misinformation. Annotated datasets are either only focused on "fake news", are small in size, or have less diversity in terms of classes.</p> <p>Here, we present a novel Twitter misinformation dataset called <strong>"CMU-MisCov19"</strong> with 4573 annotated tweets over 17 themes around the COVID-19 discourse. We also present our annotation codebook for the different COVID-19 themes on Twitter, along with their descriptions and examples, for the community to use for collecting further annotations. Further details related to the dataset, and our analysis based on this dataset can be found at <a href="https://arxiv.org/abs/2008.00791">https://arxiv.org/abs/2008.00791</a>. In adherence to the Twitter’s terms and conditions, we do not provide the full tweet JSONs but provide a ".csv" file with the tweet IDs so that the tweets can be rehydrated. We also provide the annotations, and the date of creation for each tweet for the reproduction of the results of our analyses.</p> <p><strong>Note: If for any reason, you are not able to rehydrate all the tweets, reach out to Shahan Ali Memon at (shahan@nyu.edu).</strong></p> <p>If you use this data, please cite our paper as follows: </p> <p><em>"Shahan Ali Memon and Kathleen M. Carley. Characterizing COVID-19 Misinformation Communities Using a Novel Twitter Dataset, In Proceedings of The 5th International Workshop on Mining Actionable Insights from Social Networks (MAISoN 2020), co-located with CIKM, virtual event due to COVID-19, 2020."</em></p>
MAVIS Twitter dataset: A collection of tweets and sentiment analysis in Spanish about vaccines and diseases during the period 2015-2018
<p>MAVIS dataset comprises a full knowledge base regarding Twitter messages published in Spanish during the period 2015-2018, in the context of sentiment analysis of specific vaccines and their related diseases. Such diseases and vaccines are summarized as follows:</p> <ul> <li>Invasive meningococcal disease (“EMI” in Spanish): Bexsero, Trumenba, Nimenrix</li> <li>Invasive pneumococcal disease (“ENI” in Spanish)</li> <li>Influenza</li> <li>Hepatitis</li> <li>Rotavirus: Rotarix, Rotateq</li> <li>Measles (“Sarampión” in Spanish) and MMR (“Triple vírica” in Spanish)</li> <li>Sepsis</li> <li>Whooping cough (“Tosferina” in Spanish)</li> <li>Chickenpox (“Varicela” in Spanish): Varivax, Varilrix; and Shingles (“Zoster” in Spanish)</li> <li>Human papillomavirus infection (“VPH” in Spanish): Cervarix, Gardasil</li> </ul> <p>Tweets have been manually classified as having a negative or non-negative sentiment by 5 experts. Moreover, an automatic classification has been performed by 3 different tools: IBM Watson (now Watson Tone Analyzer, <a href="https://www.ibm.com/watson/services/tone-analyzer/">https://www.ibm.com/watson/services/tone-analyzer/</a>), Google Cloud Natural Language (<a href="https://cloud.google.com/natural-language">https://cloud.google.com/natural-language</a>), and Meaning Cloud (<a href="https://www.meaningcloud.com/">https://www.meaningcloud.com/</a>). IBM Watson and Google Cloud Natural Language returned a numerical sentiment score ranging from -1 to 1, while Meaning Cloud returned a categorical variable with the values ‘P+’, ‘P’, ‘NEU’, ‘N’ and ‘N+’, which were converted to 1, 2, 3, 4 and 5 respectively.</p> <p>With these variables (IBM Watson, Google Cloud Natural Language, and Meaning Cloud annotations and the experts’ classification as the target label), a machine learning metamodel was developed. Tweets were also annotated with the sentiment output given by this classifier. </p> <p>The provided data includes intrinsic tweets information, intrinsic information regarding the users that posted the tweets, the keywords mentioned in each tweet, and the annotations that the experts, the tools, and the model gave to each tweet.</p> <p><strong>Funding</strong>: This dataset was obtained with funding from MSD, Spain under MAVIS Study (VEAP ID: 7789).</p> <p><strong>Current studies using this dataset at the moment of the publication</strong>:</p> <ul> <li>Rodríguez-González et al., “Creating a metamodel based on machine learning to identify the sentiment of vaccine and disease-related messages in Twitter: the MAVIS study” in 2020 IEEE 33st International Symposium on Computer-Based Medical Systems (CBMS), Jul. 2020, p. 6. DOI: 10.1109/CBMS49503.2020.00053</li> <li>Rodríguez-González et al., "Identifying Polarity in Tweets from an Imbalanced Dataset about Diseases and Vaccines Using a Meta-Model Based on Machine Learning Techniques" in Applied Sciences, 2020, 10. DOI: 10.3390/app10249019</li> </ul>
#IndonesiaHumanRightsSOS Twitter Hashtag Tweets Dataset
<p>Dataset ini merupakan hasil dari scraping pada media sosial twitter dengan menggunakan aplikasi twint yang ditujukan pada hashtag #IndonesiaHumanRightsSOS. Scraping data dilakukan untuk cuitan yang dibuat dari tanggal 18 Desember 2020 10:59 AM s/d 19 Desember 2020 23:18 PM.</p> <p>Pada dataset mengandung 106.903 Row data dengan informasi terkait: User ID, Username, Twitter Name,Tweets, dsb.</p> <p>Selain itu dilampirkan juga contoh data yang telah dianalisis berupa wordcloud,username cloud, 100 most used word & most active username.</p> <p>-</p> <p>This dataset is the result of scraping on social media twitter using the twint application aimed at the hashtag #IndonesiaHumanRightsSOS. Data scraping is done for tweets made from December 18 2020 10:59 AM to December 19 2020 23:18 PM.</p> <p>The dataset contains 106,903 rows of data with related information: User ID, Username, Twitter Name, Tweets, etc.</p> <p>Also there is an example of the data that has been analyzed in the form of wordcloud, username cloud, 100 most used words & most active username.</p>
Twitter analysis of the five main political leaders during the 2019 UK electoral campaign: from 12 October to 16 December 2019
<p>The analysis was conducted from 12 October to 16 December 2019 on Twitter through the study of the five main political leaders— Boris Johnson, Jeremy Corbyn, Jo Swinson, Nicola Sturgeon, and Nigel Farage — during the 2019 UK electoral campaign.</p>
Computer Scientists on Twitter
<p>This dataset contains the data used in the paper<br> <em>Identifying and Analyzing Researchers on Twitter</em> (http://dx.doi.org/10.1145/2615569.2615676).<br> At the moment, this includes computer scientists, though an extension to other disciplines is planned.<br> <br> The data can be cited as follows:<br> <em>Asmelash Teka Hadgu and Robert Jäschke. 2014. Identifying and<br> Analyzing Researchers on Twitter. In Proceedings of the 6th Annual ACM<br> Web Science Conference (WebSci '14). 23-30. ACM, New York, NY, USA. DOI: 10.1145/2615569.2615676</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.