Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
169
datasets available to search
ShareScore release 0.7.1
Dataset results
169 results for “Tweets”
Social Media Tweets Pro and Anti Bolsonaro
<p><strong>Overview</strong></p> <p>Social media platforms have an important role in Brazilian society's polarization. Especially during the COVID-19 pandemic (2020-2021), these platforms have a peak of posts about President Bolsonaro's speech and behavior during this period. On the one hand, millions of people support Bolsonaro's attitudes and follow his controversial guidances. On the other hand, a vast number of social media users accuse Bolsonaro of acting against democracy and science. </p> <p>In this context, this dataset presents the collection of tweets posts and its linked news articles (including all related media) covering two Brazilians’ demonstrations events PRO and AGAINST President Bolsonaro government, during September 7th and October 2nd of 2021.<br> The dataset contains 4.7M tweets.</p> <p><strong>Data Collection</strong><br> We use the library Fake-News Crawler (https://github.com/phillipecardenuto/fakenews-crawler) to collect the tweet posts and related media. For this, we provided keywords related to both events to receive the data during events and following days. For instance, some of the keywords used were ‘<em>7deSet</em>’, ‘<em>BolsonaroAte2026</em>’, ‘<em>VemParaRua</em>’, ‘<em>EleNao</em>’, ‘<em>07EuVou</em>’, ‘<em>SupremoÉOPovo</em>’, '<em>aculpaédobolsonaro</em>’, ‘<em>2outeuvou</em>’.<br> <br> <strong>Disclaimer</strong>: We did not perform any filtering or procedure to assert that all collected data is, in fact, related to the demonstrations; therefore, some of the content of the dataset might not be related to these events.</p> <p><strong>Content</strong><br> brazilian_demonstration_events.json: It contains the tweet posts, their metadata (e.g., post time, language), and all related media content URLs (i.e., news article link and media links).<br> <br> <strong>Media Content</strong><br> Due to the terms of use from the social networks, we do not make publicly available the images and videos that were collected. However, we can provide some extra pieces of media content related to one (or more) events by contacting the authors.<br> <br> <strong>Funding</strong><br> DéjàVu thematic project, São Paulo Research Foundation (grants 2017/12646-3, 2020/02241-9 and 2020/02211-2)<br> </p>
Corpus of political tweets UK-EU-DEBATE-20-21
<p> </p> <p>The <em>UK-EU-DEBATE-20-21</em> corpus was collected within the framework of the collaborative research project OLiNDiNUM (<em><a href="https://olindinum.huma-num.fr">Observatoire LINguistique du DIscours NUMérique</a> / </em>Linguistic Observatory of Online Debate) to be part of a shared research archive of shared corpora and resources. </p> <p>The corpus was selected with a view to examining the UK-EU media debate on the COVID-19 vaccination campaign following a specific transformative moment: the signature of the Brexit withdrawal agreement by the UK and the EU at the end of January 2021.</p> <p>The data were retrieved through the Application Programming Interface of the social networking site Twitter, using the accounts of key political actors in the UK government and EU institutions over a period of 14 months (1 February 2020–31 March 2021). The composition of the corpus is illustrated in the table.</p> <p> </p> <table> <tbody> <tr> <td><em>Political Actor</em></td> <td><em>Role</em></td> <td><em>Account</em></td> <td><em>Tweets</em></td> </tr> <tr> <td>Boris Johnson</td> <td>UK Prime Minister</td> <td>@BorisJohnson</td> <td>1186</td> </tr> <tr> <td>Dominic R. Raab</td> <td>UK Foreign Secretary</td> <td>@DominicRaab</td> <td>1468</td> </tr> <tr> <td>Priti Patel</td> <td>UK Home Secretary</td> <td>@pritipatel</td> <td>941</td> </tr> <tr> <td>Ursula von der Leyen</td> <td>President of the European Commission</td> <td>@vonderleyen</td> <td>1338</td> </tr> <tr> <td>David Sassoli</td> <td>President of the European Parliament</td> <td>@EP_President</td> <td>554</td> </tr> <tr> <td>Charles Michel</td> <td>President of the Council of the European Union</td> <td>@eucopresident</td> <td>675</td> </tr> </tbody> </table> <p> </p> <p>The data are supplied in separate .csv files (tab-delimited format). Each row contains the text of the tweet (<em>data__text</em>) and the tweet identifier (<em>data__id</em>) as a header. The tweet identifier enables swift retrieval of the original tweet by searching https://twitter.com/anyuser/status/<em>data__id. </em></p> <p> </p>
German Climate Change Tweet Corpus (GerCCT)
<p>First release of the GerCCT Corpus, a German tweet resource annotated for argument components, argument properties, sarcasm and toxic language.</p> <p>The corpus consists of 1,200 tweets and its annotations. Each tweet is associated with its respective source tweet, i.e. the tweet it replies to. Source tweets were used to provide annotators with additional context. The annotations refer to the reply tweet, i.e. NOT to the source tweet. For copyright reasons we cannot distribute the actual tweet content. Instead we share the source and reply tweet IDs and the annotations.</p> <p>The current version includes class annotations on the document level, i.e. on the tweet level. We are working on creating the respective span annotations.</p>
A Novel Dataset of Misinformation Tweets Regarding the CoronaVac Vaccine in Brazil
<p>This dataset was built to analyze the spread of misinformation about CoronaVac in Brazil by using data from Twitter for two specific events: the approval for emergency use in adults over 18 years old (January 17, 2021) and the approval for use in children aged 6 to 17 years (January 20, 2022).</p> <p>We choose to label the original tweets with at least one retweet in the analyzed period. The manual labeling of such tweets was initially performed by two annotators with high knowledge about the dataset and the considered context. In cases in which there was no agreement between the two annotators, a third annotator was considered to define the class of the tweet. </p> <p>The final dataset contains <strong>1,010 tweets from January 17, 2021</strong>, and <strong>816 tweets from January 20, 2022</strong>.</p> <p>This dataset was originally built for a conference paper accepted at BraSNAM 2022. If you make use of the dataset, please also cite the following paper:</p> <p><em>Gabriel P. Oliveira, Beatriz F. Paiva, Ana Paula Couto da Silva, and Mirella M. Moro. Characterizing the Diffusion of Misinformation Regarding the CoronaVac Vaccine in Brazil. In Proceedings of the XI Brazilian Workshop on Social Network Analysis and Mining </em><em>(BraSNAM 2022), 2022.</em></p> <pre><code>@inproceedings{brasnam/OliveiraPSM22, title = {Characterizing the Diffusion of Misinformation Regarding the CoronaVac Vaccine in Brazil}, author = {Gabriel P. Oliveira and Beatriz F. Paiva and Ana Paula Couto da Silva and Mirella M. Moro}, booktitle = {Proceedings of the XI Brazilian Workshop on Social Network Analysis and Mining (BraSNAM)} year = {2022} }</code></pre>
Tweets on Climate Central studio
<p>Tweets collected as reaction to Climate Central studio showing the sea level rise in the Catalan Coast, all tweets are in Catalan or Spanish.</p>
CCUS Sentiment Analysis - Tweets Dataset
<p>The present dataset contains Tweets in any language supported by Twitter obtained during the months January to March 2023, with any mention to the topic CCS/CCUS. The scraping process were done in Python, using the official Twitter API. All tweets were manually annotated after being machine translated into English.</p> <p><strong>- Structure </strong><br>Every row contains: <br>1st cell (A): Language <br>2nd cell (B): Tweet-text <br>3rd cell (Cc: Benefit <br>4th cell (D): Concern <br>5th cell (E): Perception – Fight climate change <br>6th cell (F): Perception – Climate-friendly technology <br>7th cell (G): Perception – Extensive R&D needed <br>8th cell (H): Perception – Better options than CCS <br>9th cell (I): Sentiment <br>10th cell (J): Relatedness <br>11th cell (K): Comments </p> <p><strong>- Annotations </strong><br><strong>Benefit </strong><br>Preventing c. change <br>Reducing c. change risks <br>Safeguarding jobs <br>Creating new jobs <br>Fossil energy production envir. friendly <br>Products envir. friendly <br>Reducing envir. impact <br>Other <br>None <br><strong>Concern </strong><br>Accidents <br>Leakages <br>Environmental <br>Earthquake-related <br>Increased local traffic <br>Investment <br>Greenwashing <br>Lock-in effects for fossil energy <br>Increase cost <br>Other <br>None <br><strong>Perception (Yes / No / None) </strong><br>Fight climate change <br>Climate-friendly technology <br>Extensive R&D needed <br>Better options than CCS <br><strong>Sentiment </strong><br>Positive <br>Negative <br>Neutral </p>
Stevia tweets
<p>This dataset includes tweets, which have been harvested from Twitter based on keywords and hashtags related to Stevia.</p>
Experimental result to investigate the influence of user's tweets and diversification on serendipitous research paper recommendations
<p>This is a raw dataset of the experiment result to investigate the influence of user's tweets and diversification on serendipitous research paper recommendations.</p> <p> </p>
Extinction Rebellion Netherlands: Dutch Climate Activism Tweets (2020-2023)
<p>This dataset contains text data from Extinction Rebellion Netherlands (XR NL) tweets between January 1, 2020, and December 31, 2023. It captures key moments in the Dutch climate activism movement, focusing on themes such as environmental protests, fossil fuel resistance, and civil disobedience. The tweets reflect XR NL’s efforts in organizing non-violent direct actions and blockades, including significant events like the A12 highway protests and Schiphol airport demonstrations. Central themes include the climate crisis, sustainability, and the ecological emergency, highlighting the movement’s focus on climate justice and the call for urgent government action in the Netherlands.</p>
Brazilian Portuguese COVID-19 Tweets
<p><strong>Brazilian Portuguese symptoms about COVID-19:</strong></p> <ul> <li><strong>Source</strong>: Twitter</li> <li><strong>Start</strong>: 2019-01-01 (January 1st)</li> <li><strong>End</strong>: 2021-09-30 (September 30th)</li> <li><strong>Tweets</strong>: 13,859,059 <ul> <li>Year 2019 [full year]: 4,043,958 obs. of 26 variables (Brazil_Portuguese_COVID19_Tweets2019.csv)</li> <li>Year 2020 [full year]: 6,155,844 obs. of 26 variables (Brazil_Portuguese_COVID19_Tweets2020.csv)</li> <li>Year 2021 [Q1 - Q3]: 3,659,257 obs. of 26 variables (Brazil_Portuguese_COVID19_Tweets2021.csv)</li> </ul> </li> </ul> <p><strong>Search terms (56 symptoms keywords about COVID-19):</strong></p> <p><strong>(1)</strong> adinamia, <strong>(2)</strong> ageusia, <strong>(3)</strong> anosmia, <strong>(4)</strong> boca azulada, <strong>(5)</strong> calafrio, <strong>(6)</strong> cansaço, <strong>(7) </strong>cefaleia, <strong>(8)</strong> cianose, <strong>(9)</strong> coloração azulada no rosto, <strong>(10)</strong> congestão nasal, <strong>(11)</strong> conjuntivite, <strong>(12) </strong>coriza, <strong>(13)</strong> desconforto respiratório, <strong>(14)</strong> diarreia, <strong>(15)</strong> dificuldade para respirar, <strong>(16)</strong> diminuição do apetite, <strong>(17)</strong> dispneia, <strong>(18)</strong> distúrbio gustativo, <strong>(19)</strong> distúrbio olfativo, <strong>(20)</strong> dor abdominal, <strong>(21)</strong> dor de cabeça, <strong>(22)</strong> dor de garganta, <strong>(23)</strong> dor no corpo, <strong>(24)</strong> dor no peito, <strong>(25) </strong>dor persistente no tórax, <strong>(26) </strong>erupção cutânea na pele, <strong>(27)</strong> fadiga, <strong>(28)</strong> falta de ar, <strong>(29)</strong> febre, <strong>(30)</strong> gripe, <strong>(31)</strong> hiporexia, <strong>(32)</strong> inapetência, <strong>(33)</strong> infecção respiratória, <strong>(34)</strong> lábio azulado, <strong>(35)</strong> mialgia, <strong>(36)</strong> nariz entupido, <strong>(37) </strong>náusea, <strong>(38)</strong> obstrução nasal, <strong>(39)</strong> perda de apetite, <strong>(40)</strong> perda do olfato, <strong>(41)</strong> perda do paladar, <strong>(42)</strong> pneumonia, <strong>(43)</strong> pressão no peito, <strong>(44)</strong> pressão no tórax, <strong>(45)</strong> prostração, <strong>(46)</strong> quadro gripal, <strong>(47)</strong> quadro respiratório, <strong>(48)</strong> queda da saturação, <strong>(49)</strong> resfriado, <strong>(50)</strong> rosto azulado, <strong>(51)</strong> saturação baixa, <strong>(52)</strong> saturação de o2 menor que 95%, <strong>(53)</strong> síndrome respiratória aguda grave, <strong>(54) </strong>srag, <strong>(55)</strong> tosse, <strong>(56)</strong> vômito.</p> <p><strong>Variables:</strong></p> <pre><code>Variable str Description ---------------------------------------------------------------------------------- id (integer64) - Tweet identifier conversation_id (integer64) - Tweet conversation identifier date (POSIXct) - Tweet created date (format: YYYY-MM-DD hh:mm:ss) tweet (chr) - Symptoms mention about COVID-19 language (chr) - Tweet language: Portuguese hashtags (chr) - Sign (#) used to identify specific topic user_id (integer64) - User identifier username (chr) - Twitter user name link (chr) - Tweet url urls (chr) - External urls from tweet photos (chr) - Photos posted in message (link) video (int) - Video posted in message (1=True;0=False) thumbnail (chr) - Thumbnail posted in message retweet (logi) - Message reposted by another user nlikes (int) - Number of tweet likes nreplies (in) - Number of tweet replies nretweets (int) - Number of tweet retweets Near (logi) - Near a certain City (Example: London) geo (logi) - Geo coordinates (lat,lon,km/mi.) user_rt_id (logi) - User retweet identifier user_rt (logi) - Retweet user retweet_id (logi) - Retweet identifier reply_to (chr) - Answer to someone retweet_date (logi) - Retweet created date (format: YYYY-MM-DD hh:mm:ss) symptoms (chr) - Symptoms mentioned nsymptoms (int) - Number of symptons mentioned </code></pre> <p><em>str: Compactly Display the Structure of an Arbitrary R Object</em></p>
MIGR-TWIT Corpus. Migration Tweets of right and far-right politics in Europe
<p><strong>Description</strong></p> <p>The <strong>MIGR-TWIT Corpus</strong> is a multilingual corpus of tweets about the topic of migration in Europe. Within the framework of the collaborative research project OLiNDiNUM (Observatoire LINguistique du DIscours NUMérique, Linguistic Observatory of Online Debate) the MIGR-TWIT Corpus is created with the aim of developing language databases of online debate. Considering the global issue of migration in line with British and French political contexts of last dozen years from 2011 to 2022, the corpus consists of two sub-corpora: </p> <ul> <li> <p><strong>FR-R-MIGR-TWIT-2011-2022 Corpus </strong>for French language data (1 January 2011 - 30 June 2022) and </p> </li> <li> <p><strong>UK-R-MIGR-RA-TWIT-2012-2022 Corpus </strong>for English language data (1 January 2012 - 5 September 2022) <strong> </strong></p> </li> </ul> <p>Using the Twitter API v2 Academic Research, tweets containing at least one occurrence of migration or refugee related words are retrieved automatically from 28 right and far-right political figures and parties. The whole corpus contains 18,233 tweets and 533,198 words. </p> <p><strong>Scientific reference:</strong></p> <p>Pietrandrea, P., Battaglia, E. (2022). “Migrants and the EU”. The diachronic construction of ad hoc categories in French far-right discourse. Journal of Pragmatics 192, 139-157.</p> <p>Blandino, G. (2023). <em>10 years of public debate on immigration: combining topic modeling and corpus linguistics to examine the British (far-)right discourse on Twitter</em>, MA University of Wolverhampton</p> <p>Jeon, S. (2025). Le discours numérique sur l'immigration en France entre 2011 et 2022. Une analyse de corpus (Online Discourse on Immigration in France between 2011 and 2022. A Corpus Analysis), PhD Thesis, Université de Lille, France.</p> <p><strong>Contents</strong></p> <p>The whole corpus contains two CSV Zip files (tabular format) corresponding to each sub-corpus. The complete corpus is presented in two versions, one version with the tweet identifier (<strong><em>data__id</em></strong>) and the text of the tweet (<strong><em>data__text</em></strong>) as a header (folders named <em>FR-R-MIGR-TWIT-2011-2022_textonly</em> and <em>UK-R-MIGR-RA-TWIT-2012-2022_textonly</em>, respectively composed of 12 and 11 Zip files of every single year), and the other version with all tweet fields information included as a header, such as the posting date (<em><strong>data__created__at</strong></em>), the username (<strong><em>author__name</em></strong>), the number of retweets (<em><strong>data__public_metrics__retweet_count</strong></em>), etc., with two folders named <em>FR-R-MIGR-TWIT-2011-2022_meta</em> and <em>UK-R-MIGR-RA-TWIT-2012-2022_meta</em>. Detailed information for each sub-corpus is illustrated below.</p> <p><strong>1. FR-R-MIGR-TWIT-2011-2022 </strong></p> <ul> <li><strong>Created at: </strong>2022-08-08</li> <li> <p><strong>Language: </strong>FR<strong> </strong></p> </li> <li> <p><strong>Coverage: </strong>16 user accounts; 11,761 tweets; 358,491 words</p> </li> <li> <p><strong>Time of data collection: </strong>start=2011-01-01; end=2022-06-30 </p> </li> <li> <p><strong>Keywords: </strong>words derived from a latin root “<em><strong>migr</strong></em>” of <em>migrare</em></p> </li> <li> <p><strong>Corpus composition: </strong></p> </li> </ul> <table> <tbody> <tr> <th> </th> <th>Political figure/party</th> <th>Username</th> <th>Tweets</th> <th>Year concerned</th> </tr> <tr> <th>1</th> <td>Michel Barnier</td> <td>@MichelBarnier</td> <td>31</td> <td>2017-22</td> </tr> <tr> <th>2</th> <td>Valérie Pécresse</td> <td>@vpecresse</td> <td>81</td> <td>2017-22</td> </tr> <tr> <th>3</th> <td>Rassemblement National</td> <td>@RNational_off</td> <td>3,347</td> <td>2017-22</td> </tr> <tr> <th>4</th> <td>Nicolas Dupont-aignan</td> <td>@dupontaignan</td> <td>663</td> <td>2011-22</td> </tr> <tr> <th>5</th> <td>Éric Ciotti</td> <td>@ECiotti</td> <td>1,007</td> <td>2012-22</td> </tr> <tr> <th>6</th> <td>Christian Estrosi</td> <td>@cestrosi</td> <td>137</td> <td>2011-22</td> </tr> <tr> <th>7</th> <td>Marine Le Pen</td> <td>@MLP_officiel</td> <td>1,650</td> <td>2011-22</td> </tr> <tr> <th>8</th> <td>Valérie Boyer</td> <td>@valerieboyer13</td> <td>837</td> <td>2012-22</td> </tr> <tr> <th>9</th> <td>Florian Philippot</td> <td>@f_philippot</td> <td>485</td> <td>2012-22</td> </tr> <tr> <th>10</th> <td>Xavier Bertrand</td> <td>@xavierbertrand</td> <td>70</td> <td>2017-22</td> </tr> <tr> <th>11</th> <td>Marion Maréchal</td> <td>@MarionMarechal</td> <td>479</td> <td>2012-17,19-22</td> </tr> <tr> <th>12</th> <td>Philippe Meunier</td> <td>@Meunier_Ph</td> <td>245</td> <td>2013-22</td> </tr> <tr> <th>13</th> <td>Jordan Bardella</td> <td>@J_Bardella</td> <td>1,095</td> <td>2013-22</td> </tr> <tr> <th>14</th> <td>Nicolas Bay</td> <td>@NicolasBay_</td> <td>1,260</td> <td>2017-22</td> </tr> <tr> <th>15</th> <td>Emmanuel Macron</td> <td>@EmmanuelMacron</td> <td>72</td> <td>2017-22</td> </tr> <tr> <th>16</th> <td>Éric Zemmour</td> <td>@ZemmourEric</td> <td>302</td> <td>2019-22</td> </tr> <tr> <th>17</th> <td>Jean Messiha*</td> <td>Banned from Twitter (since July 2021)</td> <td>-</td> <td>-</td> </tr> </tbody> </table> <ul> <li>Political figures and parties of table above are listed in chronological order according to the dates on which they posted their first tweet.</li> <li> <p><strong>*</strong>Before the launching of Twitter API v2 Academic Research, migr-tweets were collected from the database of Europresse.com including 1,453 tweets of Jean Messiha as part of the reference study (Pietrandrea & Battaglia 2022). However, the Twitter account in question has been permanently banned since July 2021. For our data collection using the Twitter API started in September 2021, we could not access this account. Therefore, we decided not to include his tweets in the FR-R-MIGR-TWIT-2011-2022 for the sake of consistency with the rest of twitter data that are automatically retrieved.</p> </li> <li> <p>The sub-corpus FR-R-MIGR-TWIT-2017-2022 is developed, annotated and analyzed as part of a doctoral thesis in progress (<a href="https://theses.fr/s360032">Jeon, 2025</a>) with the aim of studying the semantic construction of migr-lexicon over the period between 2011 and 2022. </p> </li> </ul> <p><strong> </strong></p> <p><strong>2. UK-R-MIGR-RA-TWIT-2012-2022 </strong></p> <ul> <li> <p><strong>Created at: </strong>2022-09-06</p> </li> <li> <p><strong>Language: </strong>EN</p> </li> <li> <p><strong>Coverage: </strong>12 user accounts; 6,472 tweets; 174,707 words </p> </li> <li> <p><strong>Time of data collection: </strong>start=2012-01-01; end=2022-09-05</p> </li> <li> <p><strong>Keywords: </strong>words derived from a latin root “<strong><em>migr</em></strong>” of <em>migrare </em>in addition to the keywords “<strong><em>refugee</em></strong>(<strong><em>s</em></strong>)” and “<strong><em>asylum</em></strong>”.</p> </li> <li> <p><strong>Corpus composition:</strong></p> </li> </ul> <table> <tbody> <tr> <th> </th> <th>Political figure/party</th> <th>Username</th> <th>Tweets</th> <th>Year concerned</th> </tr> </tbody> <tbody> <tr> <th>1</th> <td>David Cameron</td> <td>@David_Cameron</td> <td>32</td> <td>2012-22</td> </tr> <tr> <th>2</th> <td>Amber Rudd</td> <td>@AmberRuddUK</td> <td>29</td> <td>2012-22</td> </tr> <tr> <th>3</th> <td>Sajid Javid</td> <td>@sajidjavid</td> <td>84</td> <td>2012-22</td> </tr> <tr> <th>4</th> <td>Boris johnson</td> <td>@BorisJohnson</td> <td>80</td> <td>2015-22</td> </tr> <tr> <th>5</th> <td>Priti Patel</td> <td>@pritipatel</td> <td>304</td> <td>2012-22</td> </tr> <tr> <th>6</th> <td>UK Home Office</td> <td>@ukhomeoffice</td> <td>909</td> <td>2012-22</td> </tr> <tr> <th>7</th> <td>Nigel Farage</td> <td>@Nigel_Farage</td> <td>1,010</td> <td>2012-22</td> </tr> <tr> <th>8</th> <td>Richard Tice</td> <td>@TiceRichard</td> <td>180</td> <td>2013-22</td> </tr> <tr> <th>9</th> <td>UKIP</td> <td>@UKIP</td> <td>2,746</td> <td>2012-22</td> </tr> <tr> <th>10</th> <td>Neil Hamilton</td> <td>@NeilUKIP</td> <td>252</td> <td>2013-22</td> </tr> <tr> <th>11</th> <td>Nick Griffin</td> <td>@NickGriffinBU</td> <td>542</td> <td>2012-22</td> </tr> <tr> <th>12</th> <td>Robin Tilbrook</td> <td>@RobinTilbrook</td> <td>304</td> <td>2012-22</td> </tr> </tbody> </table> <p> </p> <ul> <li> <p>2 out of 12 accounts are official accounts belonging to the” UK Home Office” department and the “UKIP” (United Kingdom Independence Party) party. 10 out of 12 accounts are political figures’ accounts.</p> </li> <li> <p>The corpus UK-R-MIGR-RA-TWIT-2012-2022 will be exploited for the following master’s thesis: Blandino, G. (2023). <em>10 years of public debate on immigration: combining topic modeling and corpus linguistics to examine the British (far-)right discourse on Twitter</em>, MA University of Wolverhampton.</p> </li> </ul> <p> </p>
CPLP:tuítes – The pluricentric corpus of tweets in Portuguese language
<p>CPLP:tuítes is a corpus composed of 125,827 tweets and a total of 2,633,507 tokens. The tweets come from 53 newspaper accounts or news providers in Angola, Brazil, Cape Verde, Guinea-Bissau, Mozambique, Portugal, and São Tomé and Príncipe.</p>
CoVaxxy Tweet IDs data set
<p>A collection of Tweet IDs related to Covid-19 Vaccines, gathered from Twitter since Jan 4, 2021. Please see <a href="https://arxiv.org/abs/2101.07694">https://arxiv.org/abs/2101.07694</a> for more information.</p>
Multi-label Tweet Dataset for Textual Propaganda Detection related to anti-CAA protest in India (2019-2021)
<p>This is a collection of English Tweets about the Citizenship (Amendment) Bill protests that occurred in India in 2019-2021. The dataset contains tweet instances multiple labels for identified propaganda techniques. The data set consists of tweet ids, hashtags used, and corresponding propaganda techniques. Labels have been automatically generated using Weak Supervision. </p> <p>As of 2023, there are very limited textual propaganda detection dataset for Tweets. This dataset is released to facilitate future research as propaganda has become omnipresent in modern social media. </p> <p> </p> <p> </p>
MIGR-TWIT CORPORA. Migration Tweets of French Left-wing Politics.
<p><strong>Description</strong></p> <p>The <strong>FR-L-MIGR-TWIT Corpus</strong> is part of the <strong><a href="https://www.ortolang.fr/market/corpora/migr-twit-corpus">MIGR-TWIT CORPORA</a></strong>, diachronic bilingual corpus of Tweets about the topic of migration in Europe.<br>Within the framework of the collaborative research project <a href="https://olindinum.huma-num.fr/recherche/">OLiNDiNUM</a> (Observatoire LINguistique du DIscours NUMérique, [Linguistic Observatory of Online Debate]), the MIGR-TWIT Corpora are created with the aim to study the evolution of the public discourse on migration in Europe during the past dozen years from 2011 to 2022. First two components of the corpus represent migration discourse of right-wing politics in France and in the UK. The FR-L-MIGR-TWIT Corpus represents French left-wing politics' migration discourse on Twitter. </p> <p>Using the <em>Twitter API v2 Academic Research</em>, the Tweets containing at least one occurrence of lexicon derived from a latin root "<em>migr</em>" of <em>migrare </em>are automatically retrieved from 23 Twitter accounts of French left-wing political figures and parties.<br> </p> <p><strong>Scientific reference : </strong>Jeon, S. (2025). Le discours numérique sur l'immigration en France entre 2011 et 2022. Une analyse de corpus (Online Discourse on Immigration in France between 2011 and 2022. A Corpus Analysis), PhD Thesis, Université de Lille, France.</p> <p><strong>Contents</strong><br>The downloadable version of <strong>FR-L-MIGR-TWIT-2011-2022</strong> <strong>Corpus </strong>contains 32 CSV files (tabular format). The corpus is presented in simplified and complete versions in terms of metadata. The simplified version corresponds to one single file named <em><strong>FR-L-MIGR-TWIT-2011-2022.csv</strong></em>, containing four basic (meta)data, <em>i.e</em>. identifier, text, posting date and username (that is, <em><strong>data__id</strong></em>, <strong>data__text</strong>, <em><strong>data__created_at</strong></em> and <strong><em>author__name</em></strong><em> </em>as the table hearder elements). In addition to these four (meta)data, the elaborate version is provided with all Tweet fields information included as a header element, such as the numbers of Replies, Retweets, Likes and Quotes, etc. This version is also available in one single CSV file named <em><strong>FR-L-MIGR-TWIT-2011-2022_meta.csv</strong></em>.</p> <p>Besides, the elaborate version is provided with three CSV Zip files: 7 CSV files in the zip file named <em>FR-L-MIGR-TWIT-</em><strong><em>YEAR</em></strong><em>_meta</em> correspond to grouped years (<em>i.e. FR-L-MIGR-TWIT-<strong>2011-2016</strong>_meta.csv</em>) or each and every year (<em>e.g. FR-L-MIGR-TWIT-<strong>2017</strong>_meta.csv, </em>and so on) for the last dozen years. 23 files in the zip file named <em>FR-L-<strong>NAME</strong>-MIGR-TWIT_meta</em> for each and every component of selected French left-wing political figures and parties (<em>e.g. FR-L-<strong>Arthaud</strong>-TWIT_meta.csv</em>). The zip file named FR-L-MIGR-TWIT-2011-2022_meta contains yearly Tweets of each and every component of political figures and parties.</p> <p>Detailed information of the FR-L-MIGR-TWIT-2011-2022 CORPUS is illustrated below.</p> <ul> <li><strong>Created at:</strong> 2023-04-18</li> <li><strong>Language:</strong> FR</li> <li><strong>Coverage:</strong> <strong>23</strong> <strong>user accounts</strong> ; <strong>5,636 Tweets</strong> ; <strong>169,818 words</strong></li> <li><strong>Time of data collection:</strong> start=2011-01-01 ; end=2022-06-30</li> <li><strong>Keywords: </strong>words derived from a latine root “<strong><em>migr</em></strong>” of <em>migrare</em></li> <li><strong>Corpus composition:</strong></li> </ul> <table> <tbody> <tr> <td> <p> </p> </td> <td> <p><strong>Political Figure/party</strong></p> </td> <td> <p><strong>Type of representative</strong></p> </td> <td> <p><strong>Username</strong></p> </td> <td> <p><strong><em>migr</em>-Tweets</strong></p> </td> </tr> <tr> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Adrien Quatennens</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@AQuatennens</strong></p> </td> <td> <p><strong>315</strong></p> </td> </tr> <tr> <td> <p><strong>2</strong></p> </td> <td> <p><strong>Alexis Corbière</strong></p> </td> <td> <p><strong>PERSON(M)</strong></p> </td> <td> <p><strong>@Alexiscorbiere</strong></p> </td> <td> <p><strong>209</strong></p> </td> </tr> <tr> <td> <p><strong>3</strong></p> </td> <td> <p><strong>Anne Hidalgo</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@Anne_Hidalgo</strong></p> </td> <td> <p><strong>801</strong></p> </td> </tr> <tr> <td> <p><strong>4</strong></p> </td> <td> <p><strong>Arnaud Montebourg*</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@montebourg</strong></p> </td> <td> <p><strong>7</strong></p> </td> </tr> <tr> <td> <p><strong>5</strong></p> </td> <td> <p><strong>Benoît Hamon</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@benoithamon</strong></p> </td> <td> <p><strong>172</strong></p> </td> </tr> <tr> <td> <p><strong>6</strong></p> </td> <td> <p><strong>Christiane Taubira</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@ChTaubira</strong></p> </td> <td> <p><strong>11</strong></p> </td> </tr> <tr> <td> <p><strong>7</strong></p> </td> <td> <p><strong>Clémentine Autain</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@Clem_Autain</strong></p> </td> <td> <p><strong>102</strong></p> </td> </tr> <tr> <td> <p><strong>8</strong></p> </td> <td> <p><strong>Danièle Obono</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@Deputee_Obono</strong></p> </td> <td> <p><strong>415</strong></p> </td> </tr> <tr> <td> <p><strong>9</strong></p> </td> <td> <p><strong>Esther Benbassa**</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@EstherBenbassa</strong></p> </td> <td> <p><strong>936</strong></p> </td> </tr> <tr> <td> <p><strong>10</strong></p> </td> <td> <p><strong>François Hollande</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@fhollande</strong></p> </td> <td> <p><strong>28</strong></p> </td> </tr> <tr> <td> <p><strong>11</strong></p> </td> <td> <p><strong>François_Ruffin</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@Francois_Ruffin</strong></p> </td> <td> <p><strong>19</strong></p> </td> </tr> <tr> <td> <p><strong>12</strong></p> </td> <td> <p><strong>Jean-Luc Mélenchon</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@JLMelenchon</strong></p> </td> <td> <p><strong>240</strong></p> </td> </tr> <tr> <td> <p><strong>13</strong></p> </td> <td> <p><strong>Manon Aubry</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@ManonAubryFr</strong></p> </td> <td> <p><strong>182</strong></p> </td> </tr> <tr> <td> <p><strong>14</strong></p> </td> <td> <p><strong>Natalie Arthaud</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@n_arthaud</strong></p> </td> <td> <p><strong>165</strong></p> </td> </tr> <tr> <td> <p><strong>15</strong></p> </td> <td> <p><strong>Philippe Poutou</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@PhilippePoutou</strong></p> </td> <td> <p><strong>83</strong></p> </td> </tr> <tr> <td> <p><strong>16</strong></p> </td> <td> <p><strong>Raphael Glucksmann</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@rglucks1</strong></p> </td> <td> <p><strong>142</strong></p> </td> </tr> <tr> <td> <p><strong>17</strong></p> </td> <td> <p><strong>Yannick Jadot</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@yjadot</strong></p> </td> <td> <p><strong>374</strong></p> </td> </tr> <tr> <td> <p><strong>18</strong></p> </td> <td> <p><strong>Europe Écologie-Les Verts</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@EELV</strong></p> </td> <td> <p><strong>484</strong></p> </td> </tr> <tr> <td> <p><strong>19</strong></p> </td> <td> <p><strong>Gauche Républicaine et Socialiste</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@Gauche_RS</strong></p> </td> <td> <p><strong>73</strong></p> </td> </tr> <tr> <td> <p><strong>20</strong></p> </td> <td> <p><strong>Génération.s</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@GenerationsMvt</strong></p> </td> <td> <p><strong>165</strong></p> </td> </tr> <tr> <td> <p><strong>21</strong></p> </td> <td> <p><strong>La France Insoumise</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@FranceInsoumise</strong></p> </td> <td> <p><strong>300</strong></p> </td> </tr> <tr> <td> <p><strong>22</strong></p> </td> <td> <p><strong>Parti Radical Gauche</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@PartiRadicalG</strong></p> </td> <td> <p><strong>37</strong></p> </td> </tr> <tr> <td> <p><strong>23</strong></p> </td> <td> <p><strong>Parti Socialiste</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@partisocialiste</strong></p> </td> <td> <p><strong>376</strong></p> </td> </tr> </tbody> </table> <ul> <li>Political figures and parties, listed in alphabetical order, are selected according to the four criteria: (1) the high number of <em>migr</em>-tweets, (2) the political affiliation, (3) the political careers, that is, the Member of the European Parliament or (4) the presidential candidate during the period between 2011 and 2022. These four criteria are not mutually exclusive.</li> <li>As part of a doctoral thesis (<a href="https://theses.fr/s360032">Jeon, 2025</a>), the FR-L-MIGR-TWIT and FR-R-MIGR-TWIT corpora are compiled, annotated and analyzed through a comparative discourse analysis approach, with the aim to study the semantic construction of <em>migr</em>-lexicon over the period between 2011 and 2022.</li> <li>*One migration Tweet retrieved from the user account @montebourg for the year of 2019 was removed and is not included in his 7 <em>migr</em>-tweets because it refers to the issue of the migration of honey bees.</li> <li>**We later added the user account @EstherBenbassa represented by Esther Benbassa, senator and former member of political party Europe Écologie-Les Verts (representative of the user account @EELV), because of the high number of her <em>migr</em>-tweets that were retweeted by @EELV.</li> </ul> <p>The <strong>MIGR-TWIT</strong> <strong>Corpus</strong> consists of three subcorpora for a total amount of <strong>23,869 Tweets </strong>and <strong>703,016</strong> <strong>words</strong>:</p> <ul> <li>FR-R-MIGR-TWIT-2011-2022 Corpus: <em>French Right-wing</em> politics' <em>migr</em>-tweets</li> <li>UK-R-MIGR-RA-TWIT-2011-2022 Corpus: <em>British Right-wing</em> politics' <em>migr</em>-tweets</li> <li>FR-L-MIGR-TWIT-2011-2022 Corpus: <em>French Left-wing</em> politics' <em>migr</em>-tweets </li> </ul> <p> </p>
Scraped tweets about women in STEM from April 2022 to May 2023
<p>The data comprises one csv file named tweets. It has 168,677 tweets scraped with the help of snscrape spanning April 2022 to May 2023. Each entry contains metadata regarding the tweet and it's author. The tweets are curated to be representative of discourse regarding women in STEM. The search queries used while scraping are "womeninSTEM", "womeninTech", etc.</p> <p> </p>
TweetsCOV19 - A Semantically Annotated Corpus of Tweets About the COVID-19 Pandemic (Part 1, October 2019 - April 2020)
<p><strong><a href="https://data.gesis.org/tweetscov19/">TweetsCOV19</a></strong><strong> </strong>is a semantically annotated corpus of Tweets about the COVID-19 pandemic. It is a subset of <a href="https://data.gesis.org/tweetskb">TweetsKB</a> and aims at capturing online discourse about various aspects of the pandemic and its societal impact. <strong>Metadata</strong> information about the tweets as well as extracted <strong>entities</strong>, <strong>sentiments</strong>, <strong>hashtags</strong>, <strong>user mentions</strong>, and <strong>resolved URLs </strong>are exposed in RDF using established RDF/S vocabularies*.</p> <p>We also provide a <em><strong>tab-separated values (tsv)</strong></em> version of the dataset. Each line contains features of a tweet instance. Features are separated by tab character ("\t"). The following list indicate the feature indices:</p> <ol> <li>Tweet Id: Long.</li> <li>Username: String. Encrypted for privacy issues*.</li> <li>Timestamp: Format ( "EEE MMM dd HH:mm:ss Z yyyy" ).</li> <li>#Followers: Integer.</li> <li>#Friends: Integer.</li> <li>#Retweets: Integer.</li> <li>#Favorites: Integer.</li> <li>Entities: String. For each entity, we aggregated the original text, the annotated entity and the produced score from <a href="https://github.com/yahoo/FEL">FEL</a> library. Each entity is separated from another entity by char ";". Also, each entity is separated by char ":" in order to store "original_text:annotated_entity:score;". If FEL did not find any entities, we have stored "null;".</li> <li>Sentiment: String. <a href="http://sentistrength.wlv.ac.uk/">SentiStrength</a> produces a score for positive (1 to 5) and negative (-1 to -5) sentiment. We splitted these two numbers by whitespace char " ". Positive sentiment was stored first and then negative sentiment (i.e. "2 -1").</li> <li>Mentions: String. If the tweet contains mentions, we remove the char "@" and concatenate the mentions with whitespace char " ". If no mentions appear, we have stored "null;".</li> <li>Hashtags: String. If the tweet contains hashtags, we remove the char "#" and concatenate the hashtags with whitespace char " ". If no hashtags appear, we have stored "null;".</li> <li>URLs: String: If the tweet contains URLs, we concatenate the URLs using ":-: ". If no URLs appear, we have stored "null;"</li> </ol> <p>This dataset consists of <strong>8,151,524 tweets</strong> in total, posted by <strong>3,664,518 users</strong> and reflects the societal discourse about COVID-19 on Twitter in the period of October 2019 until April 2020.</p> <p>To extract the dataset from <a href="https://data.gesis.org/tweetskb">TweetsKB</a>, we compiled a seed list of 268 COVID-19-related <a href="https://data.gesis.org/tweetscov19/keywords.txt">keywords</a>.</p> <p><em>* For the sake of privacy, we anonymize user IDs and we do not provide the text of the tweets.</em></p> <p> </p>
GeoCoV19: A Dataset of Hundreds of Millions of Multilingual COVID-19 Tweets with Location Information
<p>We present GeoCoV19, a large-scale Twitter dataset related to the ongoing COVID-19 pandemic. The dataset has been collected over a period of 90 days from February 1 to May 1, 2020 and consists of more than 524 million multilingual tweets. As the geolocation information is essential for many tasks such as disease tracking and surveillance, we employed a gazetteer-based approach to extract toponyms from user location and tweet content to derive their geolocation information using the Nominatim (Open Street Maps) data at different geolocation granularity levels. In terms of geographical coverage, the dataset spans over 218 countries and 47K cities in the world. The tweets in the dataset are from more than 43 million Twitter users, including around 209K verified accounts. These users posted tweets in 62 different languages.</p>
2000 Tweets originados por la cuenta de @AlvaroUribeV
<p>2000 tweets originados en la cuenta de Alvaro Uribe. Retweets incluidos. Capturados 07/07/2020</p>
2000 Tweets de la cuenta de @AlvaroUribeV
<p>2000 tweets originados de la cuenta de Alvaro Uribe, Retweets incluidos. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.