Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
169
datasets available to search
ShareScore release 0.7.1
Dataset results
169 results for “Tweets”
Resources to compute TF-IDF weightings on press articles and tweets
<p>These two datasets of features are used in order to compute TF-IDF weightings of documents. It is meant to be used with the <a href="https://pypi.org/project/compute-tf-idf-vectors/">compute-tf-idf-vectors</a> program written in Python and available on Pypi.org.</p> <p>- features_tweets.csv contains features (tokens, lemmas and entities) extracted from Tweets published by press agencies in french, german, spanish and english.</p> <p>- features_news.csv contains features (tokens, lemmas and entities) extracted from articles published by Deutsche Welle in the same languages.</p>
Murreviikko: an Annotated and Normalized Corpus of Dialectal Finnish Tweets
<p>Murreviikko (literally 'Dialect week') is a campaign founded in the University of Eastern Finland to promote the use of Finnish dialects in social media. It started in 2020 and takes place mid-October.</p> <p>The original data was collected from Twitter with the search word murreviikko ('dialect week') and hashtag #murreviikko separately for 2020, 2021 and 2022. The current dataset combines all the original collections.</p> <p>The tweets are dialectologically annotated on two levels: following the East-West division of Finnish dialects, and following a seven-way division of Finnish dialects (South-West, Häme, Southern Ostrobothnia, Central and Northern Ostrobothnia, Far North, Savo, and South-East), appended with the Helsinki slang. There is also a class for dialectal tweets, which are not discernible (NA) because of contrasting or scarce dialectal features.</p> <p>The original tweets are normalized to a phonetic standard, but word order is not altered, or grammar rules of standard Finnish followed otherwise. This means that for instance standard Finnish possessive suffixes (minun kirja-ni 'my book-my') are not added if they are not present in the original tweet (minun kirja). Likewise, dialect words are not corrected to the standard alternative, even if such words would exist (pruukata > pruukata instead of standard tavata).</p> <p>Following the rules of the Twitter API, this repository only includes the tweet id's, dialect annotations and normalizations. The original tweets are available for scientific use by request, as granted by the European Union’s Digital Single Market directive (2019/790).</p>
COVID-19 Vaccine Tweets in Turkish
<p>This dataset contains the Turkish tweets that are collected with the keyword vaccine, sinovac and biontech in Turkish by using Twitter Academic API. The new versions will be more up-to-date and will be classified monthly folders.</p> <p>You can visit the github repository of the project:</p> <p><a href="https://github.com/burakozturan/tria-covid19">https://github.com/burakozturan/Turkish-Vaccine-Tweets</a></p> <p> </p> <p> </p> <p> </p> <p> </p>
Large Dataset of Nigeria Covid-19 Tweets for Sentiment Analysis and Opinion Mining Tasks
<p><strong>Background</strong></p> <p>Information is essential for growth; without it, little can be accomplished. Data gathering has seen significant changes throughout the previous few centuries because of certain transitory medium. The look and style of information transference are affected by the employment of new and emerging technologies, some of which are efficient, others are reliable, and many more are quick and effective, but a few were disappointing for various reasons.</p> <p><strong>Aims</strong></p> <p>This study aims at using TextBlob and VADER analyser with historical tweets, to analyse emotional responses to the corona virus pandemic (covid-19). It shows us how much of a sociological, environmental, and economic impact it has in Nigeria, among other things. This study would be a tremendous step forward for students, researchers, and scholars who want to advance in fields like data science, machine learning, and deep learning.</p> <p><strong>Methodology</strong></p> <p>The hashtag ‘covid-19' was used to collect 1,048,575 tweets from Twitter. The tweets were pre-processed with a twitter tokenizer, and Valence Aware Dictionary for Sentiment Reasoning (VADER) and TextBlob were used for sentiment and text mining, respectively. Topic modelling was done with Latent Dirichlet Allocation (LDA). The simulated subjects, on the other hand, were visualized using Multidimensional scaling (MDS).</p> <p><strong>Results</strong></p> <p>The result of the VADER sentiment returned 39.8%, 31.3% and 28.9%, positive, neutral, and negative sentiment respectively while the result of the TextBlob sentiment returned 46.0%, 36.7% and 17.3%, neutral, positive, and negative sentiment, respectively.</p> <p><strong>Conclusion</strong></p> <p>With all of this, information from social media may be used to help organizations, governments, and nations around the world make smart and effective decisions about how to restrict and limit the negative effects of covid-19. Also know the opinion and challenges of people, then deal with problem of misinformation.</p> <p>It is concluded that with popular belief a significant number of the populace regards covid-19 as a virus that has come to stay, some believe it will eventually be conquered.</p>
Japanese "Olympic" and "Suga" Tweets from 2021-07-01 to 2021-09-02
<p><strong>Abstract</strong> (our paper)</p> <p>The Olympic Games are a typical media event and are seen as a festive occasion that monopolizes people’s attention through the mass media. The Games and their media coverage have a predetermined schedule that enhances the nation’s sense of unity by placing a temporary truce on political conflicts. Governments, especially those of Olympic host countries, tend to take advantage of this effect to garner support for their own policies. However, the effects of such media events may be weakening owing to changes in the media environment and increasing political polarization. Examining the 2020 Tokyo Olympic Games, this case study analyzes a large amount of Twitter data to probe Japanese social media users’ attitudes toward the Olympic Games and the relationship of these attitudinal changes with their attitudes toward the political leadership of the prime minister. The results showed that previously negative attitudes toward the Olympic Games improved as people enjoyed the event. However, this positive shift did not appear to be associated with their attitudes toward the prime minister. Users’ political predispositions strongly determined their attitudes toward the Olympic Games, indicating that the Olympic Games as a media event had limited implications for support for the administration.</p> <p><strong>Data</strong></p> <p>olympic.tsv.gz, suga.tsv.gz:<br> The first column is the tweet id, the second column is the tweet id of the retweet source, and the third column is the date and time (JST) when the tweet was posted.<br> This data was collected by giving the query "オリンピック OR パラリンピック OR 五輪 OR 聖火 OR オリパラ" or "スガ OR 菅 OR 総理 OR 首相" to the Twitter Search API. Therefore, most of the tweets are Japanese tweets. The second column is empty if the tweet is not a retweet.</p> <p><strong>Publication</strong></p> <p>This data set was created for our study. If you make use of this data set, please cite:<br> Takeshi Sakaki, Tetsuro Kobayashi, Mitsuo Yoshida, Fujio Toriumi. Do media events still unite the host nation’s citizens? The case of the Tokyo 2020 Olympic Games. <em>PLOS ONE</em>. vol.17, no.12, e0278911, 2022.<br> <a href="https://doi.org/10.1371/journal.pone.0278911">https://doi.org/10.1371/journal.pone.0278911</a></p>
Emergency management/Natural Hazards: annotated tweets
<p>A set of annotated tweets related to natural hazards and emergency management.</p> <p>Information available for each tweet:</p> <p>- tweet id</p> <p>- boolean flags about its content: floods;storms;landslides;snow;infrastructures;affected individuals;caution advice;donations & volunteering;emotional support;other info;panic</p> <p> </p>
Japanese "Abe" Tweets from 2019-02-10 to 2020-10-07 (15,407,811 tweets and 114,231,250 retweets)
<p><strong>Abstract</strong> (our paper)</p> <p>To examine conservative–liberal differences in the extent to which partisan tweets reach less partisan moderate users in a nonwestern context, we analyzed a network of retweets about former Japanese Prime Minister Shinzo Abe. The analyses consistently demonstrated that partisan tweets originating from the conservative cluster reach a wider range of moderate users than those from the liberal cluster. Network analyses revealed that while the conservative and the liberal clusters’ internal structures were similar, the conservative cluster reciprocated the follows from moderate accounts at a higher rate than the liberal cluster. In addition, moderate accounts reciprocated the conservative cluster’s following at a higher rate than they did for the liberal cluster. The analysis of tweet content showed no difference in the frequency of hashtag use between conservatives and liberals, but there were differences in the use of emotion words and linguistic expressions. In particular, emotion words related to the propagation of messages, such as those expressing “dislike”, were used more frequently by conservatives, while the use of adjectives by conservatives was closer to that of moderate users, indicating that conservative tweets are more palatable for moderate users than liberal tweets.</p> <p><strong>Data</strong></p> <p>Abe.tsv.gz:<br> The first column is the tweet id, the second column is the tweet id of the retweet source, and the third column is the date and time (JST) when the tweet was posted.<br> This data was collected by giving the query "安倍 OR アベ" to the Twitter Search API. Therefore, most of the tweets are Japanese tweets. The second column is empty if the tweet is not a retweet.</p> <p><strong>Publication</strong></p> <p>This data set was created for our study. If you make use of this data set, please cite:<br> Mitsuo Yoshida, Takeshi Sakaki, Tetsuro Kobayashi, Fujio Toriumi. Japanese conservative messages propagate to moderate users better than their liberal counterparts on Twitter. <em>Scientific Reports</em>. vol.11, article no.19224, 2021.<br> <a href="https://doi.org/10.1038/s41598-021-98349-2">https://doi.org/10.1038/s41598-021-98349-2</a></p>
French Entity-Linking dataset between annotated tweets collected during major crises in France and French Wikipedia corpus
<p>Most of the available datasets are not particularly adapted to our target application: geolocate natural disasters from social networks. First, social media posts are largely underrepresented in these datasets, and the only Twitter dataset lacks Entity-Linking annotations. Second, none of the datasets focuses on a crisis or natural disaster event.</p> <p>To mitigate these issues, we extracted a collection of French tweets written during earthquakes and major floods that have occurred in France in recent years. We set up Label-Studio in order to annotate these tweets. A total of 4617 tweets were annotated, including 1678 tweets posted during earthquakes and 2939 during floods. For each annotated tweet, mentions were annotated using the set of labels described earlier in the paper as well as, when possible, the target Wikipedia title.</p> <p>Named “RéSoCIO” in reference to the research project in which it was carried out, the dataset resulting from this work contains a total of 12 828 annotated mentions and 1 513 distinct Wikipedia entities. 85% of mentions were associated with a Wikipedia page and 94 % if we ignore the RISKNAT and DAMAGES labels, which are often difficult to map to an existing entity.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entities</strong></td> </tr> <tr> <td>PERSON</td> <td>315</td> <td>263</td> <td>136</td> </tr> <tr> <td>ORG</td> <td>863</td> <td>790</td> <td>281</td> </tr> <tr> <td>GEOLOC</td> <td>4375</td> <td>4234</td> <td>701</td> </tr> <tr> <td>TRANSPORT</td> <td>250</td> <td>203</td> <td>101</td> </tr> <tr> <td>EVENT</td> <td>35</td> <td>21</td> <td>16</td> </tr> <tr> <td>FACILITY</td> <td>129</td> <td>94</td> <td>49</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>128</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>223</td> <td>200</td> <td>46</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>12828</strong></td> <td><strong>1322</strong></td> <td><strong>1513</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the Twitter dataset. #Mentions shows the total number of mentions per label, #Linked the number of mentions linked to an entity and #Entities the number of distinct entities per label present in the dataset.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entitie</strong>s</td> </tr> <tr> <td>PERSON</td> <td>1100102</td> <td>1098406</td> <td>557697</td> </tr> <tr> <td>ORG</td> <td>750925</td> <td>749504</td> <td>130394</td> </tr> <tr> <td>GEOLOC</td> <td>2729702</td> <td>2728296</td> <td>215924</td> </tr> <tr> <td>TRANSPORT</td> <td>161539</td> <td>160487</td> <td>53405</td> </tr> <tr> <td>EVENT</td> <td>798433</td> <td>798251</td> <td>86471</td> </tr> <tr> <td>FACILITY</td> <td>258835</td> <td>258513</td> <td>109867</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>127</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>4340621</td> <td>4339658</td> <td>682458</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>10146795</strong></td> <td><strong>10138230</strong></td> <td><strong>1836399</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the full dataset. #Mentions shows the total number of mentions per label, #Linked the number of mentions linked to an entity and #Entities the number of distinct entities per label present in the dataset.</p>
Tweets containing "climate change" with topic annotations
<p>This dataset contains the Twitter IDs of all ~20M tweets containing the phrase "climate change" 2018-2021. Additionally, it contains the topical annotations and 2D semantic representation of our thematic analysis based on ~980 topic clusters that are grouped by hand into seven themes (COVID-19, Politics, Contrarian, Movements, Solutions, Impacts, Causes) as well as "non-relevant/spam", "others", and highlighting of potentially interesting topics.</p> <p>Code and additional notes are available on GitHub: https://github.com/TimRepke/twitter-climate</p> <p>The topics, including statistics and the annotator labels for broader themes (aka "super topics") are contained in the spreadsheet. This data is extrapolated to the tweets contained in the share.jsonl file containing one json object per line with the following fields:</p> <ul> <li><strong>'rel':</strong> true iff Tweet is contained in analysis</li> <li><strong>'filters':</strong> null if Tweet is not included, otherwise contains an object with "reasons" why this tweet was excluded <ul> <li><strong>'dup'</strong>: 1 iff this is a duplicate (excl first)</li> <li><strong>'lan':</strong> 1 iff language is English (and not None)</li> <li><strong>'txt': </strong>1 iff status text is not None</li> <li><strong>'mit': </strong>1 iff text has minimum number of tokens (>=4)</li> <li><strong>'mah'</strong>: 1 iff text has less than maximum number of hashtags (<=5),</li> <li><strong>'pfd'</strong>: 1 iff tweet was posted after 01.01.2018</li> <li><strong>'ptd'</strong>: 1 iff tweet was posted before 31.12.2021</li> <li><strong>'cli':</strong> 1 iff tweet actually contains "climate change" (API matches some false positives)</li> </ul> </li> <li><strong>'ann'</strong>: null if Tweet is not included, otherwise contains an object with topic annotations <ul> <li><strong>'t_km': </strong> topic (based on "keep & majority vote" strategy)</li> <li><strong>'t_kp': </strong> topic (based on "keep & closest topic centroid [proximity]" strategy)</li> <li><strong>'t_fm': </strong> topic (based on "drop sample topic [fresh] & majority vote" strategy)</li> <li><strong>'t_fp':</strong> topic (based on "drop sample topic [fresh] & closest topic centroid [proximity]")</li> <li><strong>'st_int':</strong> theme annotation "Interesting"</li> <li><strong>'st_nr': </strong> theme annotation "Non-relevant / spam"</li> <li><strong>'st_cov':</strong> theme annotation "COVID"</li> <li><strong>'st_pol': </strong> theme annotation "Politics"</li> <li><strong>'st_mov': </strong> theme annotation "Movements"</li> <li><strong>'st_imp':</strong> theme annotation "Impacts"</li> <li><strong>'st_cau': </strong> theme annotation "Causes"</li> <li><strong>'st_sol': </strong> theme annotation "Solutions"</li> <li><strong>'st_con': </strong>theme annotation "Contrarian"</li> <li><strong>'st_oth': </strong> theme annotation "Other"</li> <li><strong>'x': </strong> x position in 2D representation</li> <li><strong>'y': </strong> x position in 2D representation</li> <li><strong>'sample':</strong> true iff this tweet was in the original topic model sample</li> </ul> </li> </ul>
COVID-19 Tweets : A dataset contaning more than 600k tweets on the novel CoronaVirus
<p>This dataset contains 653 996 tweets related to the Coronavirus topic and highlighted by hashtags such as: #COVID-19, #COVID19, #COVID, #Coronavirus, #NCoV and #Corona. The tweets' crawling period started on the 27<sup>th</sup> of February and ended on the 25<sup>th</sup> of March 2020, which is spread over four weeks. </p> <p>The tweets were generated by 390 458 users from 133 different countries and were written in 61 languages. English being the most used language with almost 400k tweets, followed by Spanish with around 80k tweets. </p> <p>The data is stored in as a CSV file, where each line represents a tweet. The CSV file provides information on the following fields:</p> <ul> <li>Author: the user who posted the tweet</li> <li>Recipient: contains the name of the user in case of a reply, otherwise it would have the same value as the previous field</li> <li>Tweet: the full content of the tweet</li> <li>Hashtags: the list of hashtags present in the tweet</li> <li>Language: the language of the tweet</li> <li>Relationship: gives information on the type of the tweet, whether it is a retweet, a reply, a tweet with a mention, etc. </li> <li>Location: the country of the author of the tweet, which is unfortunately not always available</li> <li>Date: the publication date of the tweet</li> <li>Source: the device or platform used to send the tweet</li> </ul> <p>The dataset can as well be used to construct a social graph since it includes the relations "Replies to", "Retweet", "MentionsInRetweet" and "Mentions".</p>
MAVIS Twitter dataset: A collection of tweets and sentiment analysis in Spanish about vaccines and diseases during the period 2015-2018
<p>MAVIS dataset comprises a full knowledge base regarding Twitter messages published in Spanish during the period 2015-2018, in the context of sentiment analysis of specific vaccines and their related diseases. Such diseases and vaccines are summarized as follows:</p> <ul> <li>Invasive meningococcal disease (“EMI” in Spanish): Bexsero, Trumenba, Nimenrix</li> <li>Invasive pneumococcal disease (“ENI” in Spanish)</li> <li>Influenza</li> <li>Hepatitis</li> <li>Rotavirus: Rotarix, Rotateq</li> <li>Measles (“Sarampión” in Spanish) and MMR (“Triple vírica” in Spanish)</li> <li>Sepsis</li> <li>Whooping cough (“Tosferina” in Spanish)</li> <li>Chickenpox (“Varicela” in Spanish): Varivax, Varilrix; and Shingles (“Zoster” in Spanish)</li> <li>Human papillomavirus infection (“VPH” in Spanish): Cervarix, Gardasil</li> </ul> <p>Tweets have been manually classified as having a negative or non-negative sentiment by 5 experts. Moreover, an automatic classification has been performed by 3 different tools: IBM Watson (now Watson Tone Analyzer, <a href="https://www.ibm.com/watson/services/tone-analyzer/">https://www.ibm.com/watson/services/tone-analyzer/</a>), Google Cloud Natural Language (<a href="https://cloud.google.com/natural-language">https://cloud.google.com/natural-language</a>), and Meaning Cloud (<a href="https://www.meaningcloud.com/">https://www.meaningcloud.com/</a>). IBM Watson and Google Cloud Natural Language returned a numerical sentiment score ranging from -1 to 1, while Meaning Cloud returned a categorical variable with the values ‘P+’, ‘P’, ‘NEU’, ‘N’ and ‘N+’, which were converted to 1, 2, 3, 4 and 5 respectively.</p> <p>With these variables (IBM Watson, Google Cloud Natural Language, and Meaning Cloud annotations and the experts’ classification as the target label), a machine learning metamodel was developed. Tweets were also annotated with the sentiment output given by this classifier. </p> <p>The provided data includes intrinsic tweets information, intrinsic information regarding the users that posted the tweets, the keywords mentioned in each tweet, and the annotations that the experts, the tools, and the model gave to each tweet.</p> <p><strong>Funding</strong>: This dataset was obtained with funding from MSD, Spain under MAVIS Study (VEAP ID: 7789).</p> <p><strong>Current studies using this dataset at the moment of the publication</strong>:</p> <ul> <li>Rodríguez-González et al., “Creating a metamodel based on machine learning to identify the sentiment of vaccine and disease-related messages in Twitter: the MAVIS study” in 2020 IEEE 33st International Symposium on Computer-Based Medical Systems (CBMS), Jul. 2020, p. 6. DOI: 10.1109/CBMS49503.2020.00053</li> <li>Rodríguez-González et al., "Identifying Polarity in Tweets from an Imbalanced Dataset about Diseases and Vaccines Using a Meta-Model Based on Machine Learning Techniques" in Applied Sciences, 2020, 10. DOI: 10.3390/app10249019</li> </ul>
#IndonesiaHumanRightsSOS Twitter Hashtag Tweets Dataset
<p>Dataset ini merupakan hasil dari scraping pada media sosial twitter dengan menggunakan aplikasi twint yang ditujukan pada hashtag #IndonesiaHumanRightsSOS. Scraping data dilakukan untuk cuitan yang dibuat dari tanggal 18 Desember 2020 10:59 AM s/d 19 Desember 2020 23:18 PM.</p> <p>Pada dataset mengandung 106.903 Row data dengan informasi terkait: User ID, Username, Twitter Name,Tweets, dsb.</p> <p>Selain itu dilampirkan juga contoh data yang telah dianalisis berupa wordcloud,username cloud, 100 most used word & most active username.</p> <p>-</p> <p>This dataset is the result of scraping on social media twitter using the twint application aimed at the hashtag #IndonesiaHumanRightsSOS. Data scraping is done for tweets made from December 18 2020 10:59 AM to December 19 2020 23:18 PM.</p> <p>The dataset contains 106,903 rows of data with related information: User ID, Username, Twitter Name, Tweets, etc.</p> <p>Also there is an example of the data that has been analyzed in the form of wordcloud, username cloud, 100 most used words & most active username.</p>
URLs from tweets for a 2014 sample of Twitter users and for a set of computer scientists
<p>The files in this dataset are used to analyse the tweeting behaviour of computer scientists on Twitter. They comprise</p> <ul> <li>a set of 989,529 tweet-URL pairs (<em>tweets_2014_researcher.tsv.bz2</em>) from 2014 from 6,271 users of the computer scientists sample in https://zenodo.org/record/12942 specified by time, tweet id, user id, and URL,</li> <li>a set of 300,053,850 tweet ids (<em>tweets_2014_sample.tsv.bz2</em>) from the 1% Twitter stream sample from 2014,</li> <li>a set of 671,304 tweet-URL pairs (<em>tweets_2014_sample_6271_users.tsv.bz2</em>) from the 1% Twitter stream sample from 2014 for 6,271 users specified by time, tweet id, user id, and URL,</li> <li>a set of the top 10,000 host names (<em>MAG_hosts_10000.tsv</em>) from the Microsoft Academic Graph data (http://blogs.msdn.com/b/msr_er/archive/2015/06/26/announcing-the-microsoft-academic-graph-let-the-research-begin.aspx), specified by rank, URL count, and host name, and</li> <li>a set of 340 host names of URL shortening services (<em>url_shortening_services.tsv</em>).</li> </ul>
URLs from tweets for a 2014 sample of Twitter users and for a set of computer scientists
<p>The files in this dataset are used to analyse the tweeting behaviour of computer scientists on Twitter. They comprise</p> <ul> <li>a set of 989,529 tweet-URL pairs (<em>tweets_2014_researcher.tsv.bz2</em>) from 2014 from 6,271 users of the computer scientists sample in https://zenodo.org/record/12942 specified by time, tweet id, user id, and URL,</li> <li>a set of 300,053,850 tweet ids (<em>tweets_2014_sample.tsv.bz2</em>) from the 1% Twitter stream sample from 2014,</li> <li>a set of 605,080 tweet-URL pairs (<em>tweets_2014_sample_6694_users.tsv.bz2</em>) from the 1% Twitter stream sample from 2014 for 6,694 users specified by time, tweet id, user id, and URL,</li> <li>a set of the top 10,000 host names (<em>MAG_hosts_10000.tsv</em>) from the Microsoft Academic Graph data (http://blogs.msdn.com/b/msr_er/archive/2015/06/26/announcing-the-microsoft-academic-graph-let-the-research-begin.aspx), specified by rank, URL count, and host name, and</li> <li>a set of 340 host names of URL shortening services (<em>url_shortening_services.tsv</em>).</li> </ul> <p>In addition, the following rankings (based on the odds ratio) of domains, hosts, and URLs that appear in both the researcher dataset and the sample are included:</p> <ul> <li><em>domains_by_odds_ratio.tsv.bz2</em> - a ranking of 61,860 domains,</li> <li><em>hosts_by_odds_ratio.tsv.bz2</em> - a ranking of 80,384 hosts,</li> <li><em>publisher_domains_by_odds_ratio.tsv.bz2</em> - a ranking of 924 publisher domains,</li> <li><em>publisher_urls_by_odds_ratio.tsv.bz2</em> - a ranking of 4,227 publisher URLs.</li> </ul>
Toward a Comparable Corpus of Latvian, Russian and English Tweets
<p>Twitter has become a rich source for linguistic data. Here, a possibility of building a trilingual Latvian-Russian-English corpus of tweets from Riga, Latvia is investigated. Such a corpus, once constructed, might be of great use for multiple purposes such as training machine translation models, examining cross-lingual phenomena and studying the population of Riga. This pilot study shows that it is feasible to build such a resource by building and analysing a pilot corpus, which is made publicly available and can be used to construct a large comparable corpus.</p>
Sample tweets from May 2022
<p>Dataset containing twitter data, namely hashed twitter id, hashed user id, tweet language, user statistics.</p>
A German Language Labeled Dataset of Tweets
<p>Our dataset contains 8,048 German language tweets related to Jewish life from a four-year timespan. </p><p>The dataset consists of 18 samples of tweets with the keyword "Juden" or "Israel." The samples are representative samples of all live tweets (at the time of sampling) with these keywords respectively over the indicated time period. Each sample was annotated by two expert annotators using an Annotation Portal that visualizes the live tweets in context. We provide the annotation results based on the agreement of two annotators, after discussing discrepancies (Jikeli et al. 2022: 3-6). </p><p> Overall, 335 tweets (4%) were labelled as antisemitic following the IHRA Working Definition of Antisemitism. 1345 tweets (17 %) come from 2019, 1364 tweets (17 %) from 2020, 2639 tweets (33 %) from 2021 and 2700 tweets (34 %) from 2022. </p><p>About half of the tweets, a total of 4,493 tweets (56 %) come from queries with the keyword "Juden," which is representative of a continuous time period from January 2019 to December 2022: 864 tweets (19 %) come from 2019, 891 tweets (20 %) from 2020, 1364 tweets (30 %) from 2021 and 1374 (31 %). 148 out of the 4493 tweets, so 3% from the query with "Juden" are antisemitic. </p><p>The other part of the tweets, a total of 3,555 (44 %) results of queries with the keyword "Israel". 481 tweets (14 %) of the keywords containing Israel stem from 2019, 473 (13 %) come from 2020, 1275 tweets (36 %) from 2021 and 1326 tweets (37 %) are from 2022. Out of all tweets from the "Israel" query, 187 (5 %) are antisemitic. </p><p>The csv file contains diacritics and special characters of the German language (e.g., "ä", "ü", "ö", "ß"), which should be taken into account when opening it with anything other than a text editor. </p><p><strong>Acknowledgements</strong> </p><p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296. </p><p>We are grateful for the support of Indiana University's Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media & Hate Research Lab at Indiana University's Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenhäusler, Sophie von Máriássy, Mabel Poindexter, Jenna Solomon, Clara Schilling, Emma Shriberg and Victor Tschiskale. </p>
Tweets on carbon dioxide removal
<div> <h1><span>Growing online attention and positive sentiments towards</span><br><span>carbon dioxide removal</span></h1> <p>Tim Repke, Finn Müller-Hansen, Emily Cox, Jan Minx</p> <p><em><span>(unpublished)</span></em></p> <p><span>Scaling up CO2 removal is crucial to achieve net-zero targets and limit global warming. </span><span>Understanding public perception of large-scale carbon dioxide removal (CDR) is vital to </span><span>avoid opposition that could slow down development, investments, and deployment. <br></span><span>Using Twitter data from 2010 to 2022, we analysed attention and sentiments towards ten </span><span>CDR methods. Our study provides up-to-date time series evidence complementing survey </span><span>studies, capturing the opinions of users with knowledge or awareness of emerging CDR </span><span>methods. </span><span>Attention towards CDR has grown exponentially, particularly in recent years. <br></span><span>Overall, </span><span>the discourse on CDR has become more positive, except for BECCS. Conventional CDR </span><span>methods are the most discussed and receive more positive sentiments. We examined three </span><span>user types, each with varying levels of involvement in the discourse.</span> <span>Infrequent users (as</span><span>sumed less familiar) pay more attention to methods with biological sinks, while frequent </span><span>users (assumed more familiar) focus more on novel CDR methods.</span></p> <h3><span>Dataset description</span></h3> <p><strong><span>export.csv<br></span></strong><span>This dataset contains all Twitter IDs of tweets retrieved for this study. We also include all technology annotations, which technology-specific subquery this relates to, sentiment classification, and user type. Note, that this dataset contains more tweets than used in the study.</span></p> <p><strong>UserPanels.csv<br></strong>This is an overview of all users, their categorisation, and the number of positive/negative/neutral tweets. This could also be derived from the export.csv.</p> </div> <div> <div> </div> <div> </div> </div>
IA Tweets Analysis Dataset (Spanish)
<h3><strong>Cite as</strong></h3> <p><em><strong>Guerrero-Contreras, G., Balderas-Díaz, S., Serrano-Fernández, A., & Muñoz, A. (2024, June). Enhancing Sentiment Analysis on Social Media: Integrating Text and Metadata for Refined Insights. In 2024 International Conference on Intelligent Environments (IE) (pp. 62-69). IEEE.</strong></em></p> <h3>General Description</h3> <p>This dataset comprises 4,038 tweets in Spanish, related to discussions about artificial intelligence (AI), and was created and utilized in the publication "Enhancing Sentiment Analysis on Social Media: Integrating Text and Metadata for Refined Insights," (<a href="https://doi.org/10.1109/IE61493.2024.10599899" target="_blank" rel="noopener">10.1109/IE61493.2024.10599899</a>) presented at the 20th International Conference on Intelligent Environments. It is designed to support research on public perception, sentiment, and engagement with AI topics on social media from a Spanish-speaking perspective. Each entry includes detailed annotations covering sentiment analysis, user engagement metrics, and user profile characteristics, among others.</p> <h3>Data Collection Method</h3> <p>Tweets were gathered through the Twitter API v1.1 by targeting keywords and hashtags associated with artificial intelligence, focusing specifically on content in Spanish. The dataset captures a wide array of discussions, offering a holistic view of the Spanish-speaking public's sentiment towards AI.</p> <h3>Dataset Content</h3> <ul> <li><strong>ID</strong>: A unique identifier for each tweet.</li> <li><strong>text</strong>: The textual content of the tweet. It is a string with a maximum allowed length of 280 characters.</li> <li><strong>polarity</strong>: The tweet's sentiment polarity (e.g., Positive, Negative, Neutral).</li> <li><strong>favorite_count</strong>: Indicates how many times the tweet has been liked by Twitter users. It is a non-negative integer.</li> <li><strong>retweet_count</strong>: The number of times this tweet has been retweeted. It is a non-negative integer.</li> <li><strong>user_verified</strong>: When true, indicates that the user has a verified account, which helps the public recognize the authenticity of accounts of public interest. It is a boolean data type with two allowed values: True or False.</li> <li><strong>user_default_profile</strong>: When true, indicates that the user has not altered the theme or background of their user profile. It is a boolean data type with two allowed values: True or False.</li> <li><strong>user_has_extended_profile</strong>: When true, indicates that the user has an extended profile. An extended profile on Twitter allows users to provide more detailed information about themselves, such as an extended biography, a header image, details about their location, website, and other additional data. It is a boolean data type with two allowed values: True or False.</li> <li><strong>user_followers_count</strong>: The current number of followers the account has. It is a non-negative integer.</li> <li><strong>user_friends_count</strong>: The number of users that the account is following. It is a non-negative integer.</li> <li><strong>user_favourites_count</strong>: The number of tweets this user has liked since the account was created. It is a non-negative integer.</li> <li><strong>user_statuses_count</strong>: The number of tweets (including retweets) posted by the user. It is a non-negative integer.</li> <li><strong>user_protected</strong>: When true, indicates that this user has chosen to protect their tweets, meaning their tweets are not publicly visible without their permission. It is a boolean data type with two allowed values: True or False.</li> <li><strong>user_is_translator</strong>: When true, indicates that the user posting the tweet is a verified translator on Twitter. This means they have been recognized and validated by the platform as translators of content in different languages. It is a boolean data type with two allowed values: True or False.</li> </ul> <h3>Potential Use Cases</h3> <p>This dataset is aimed at academic researchers and practitioners with interests in:</p> <ul> <li>Sentiment analysis and natural language processing (NLP) with a focus on AI discussions in the Spanish language.</li> <li>Social media analysis on public engagement and perception of artificial intelligence among Spanish speakers.</li> <li>Exploring correlations between user engagement metrics and sentiment in discussions about AI.</li> </ul> <h3>Data Format and File Type</h3> <p>The dataset is provided in CSV format, ensuring compatibility with a wide range of data analysis tools and programming environments.</p> <h3>License</h3> <p>The dataset is available under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting sharing, copying, distribution, transmission, and adaptation of the work for any purpose, including commercial, provided proper attribution is given.</p>
AWOFRO : Annotated tweet corpus of mixed Wolof-French for detecting obnoxious messages
<p>These data are tweets of mixed Wolof-French codes annotated by three(3) annotators. <br>They were extracted during the period from 1 January 2021 to 31 May 2023.</p> <p>Content description :</p> <p><strong>Corpora.rar</strong> : The dataset contains 3510 annotated tweets</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.