Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
169
datasets available to search
ShareScore release 0.7.1
Dataset results
169 results for “Tweets”
Datasets of ASONAM-2015 paper "Tweet sentiment: From classification to quantification"
<p>Datasets used for the following ASONAM 2015 paper:<br> ---------------------------------------------------------------------------------------------------<br> Title: Tweet Sentiment: From Classification to Quantification<br> Authors: Wei Gao and Fabrizio Sebastiani<br> Organization: Qatar Computing Research Institute, Hamad Bin Khalifa University, Doha, Qatar<br> ---------------------------------------------------------------------------------------------------</p> <p>[Content]</p> <p>* SemEval2013, SemEval2014, SemEval2015 datasets:<br> - semeval.train.feature.txt: Training set for learning sentiment models at development stage<br> - semeval.dev.feature.txt: Held-out set for tuning parameters<br> - semeval.train+dev.feature.txt: Training set for learning the final sentiment model<br> - semeval13.test.feature.txt: SemEval2013 test set<br> - semeval14.test.feature.txt: SemEval2014 test set<br> - semeval15.test.feature.txt: SemEval2015 test set<br> <br> * Other datasets: sanders, sst, omd, hcr, gasp<br> - X.train.feature.txt: Training set for learning sentiment models at development stage<br> - X.dev.feature.txt: Held-out set for tuning parameters<br> - X.train+dev.feature.txt: Traing set for learning the final sentiment model<br> - X.test.feature.txt: Test set<br> where X is one of sanders, sst, omd, hcr and gasp.</p> <p>For more details, please refer to the paper.</p> <p><br> [Citation]<br> You can cite the folowing paper when referring to the dataset:</p> <p>@inproceedings{gao2015tweet,<br> title={Tweet sentiment: From classification to quantification},<br> author={Gao, Wei and Sebastiani, Fabrizio},<br> booktitle={2015 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM)},<br> pages={97--104},<br> year={2015},<br> organization={IEEE}<br> }</p>
A Dataset of State-Censored Tweets
<p>This is the dataset associated with the paper of the same name. You can find it here: <a href="https://arxiv.org/abs/2101.05919">https://arxiv.org/abs/2101.05919</a></p> <p><br> <strong>Files:</strong></p> <ul> <li>tweets.csv : All 583k censored tweets</li> <li>tweets_debiased.csv : Debiased sample of tweets (Section 6.1)</li> <li>all_users.csv : All users who are censored once at least once</li> <li>users.csv : All 4301 users whose entire profile is censored</li> <li>users_inferred.csv : 1931 extra users inferred to be the censored by the procedure described in Section 3.3</li> <li>supplement.csv : The supplementary tweet data. (Section 3.5)</li> </ul> <p>Please refer to this Github repo for the detailed documentation and the code for reproduction.</p> <p><a href="https://github.com/tugrulz/CensoredTweets">https://github.com/tugrulz/CensoredTweets</a></p>
#paris #Bataclan #parisattacks #porteouverte tweets
<p>Tweet ids for #paris #Bataclan #parisattacks #porteouverte tweets. Tweets can be "hydrated" with Ed Summers' twarc (https://github.com/edsu/twarc). twarc.py --hydrate paris-tweet-ids.txt > paris-tweets.json. Hydrating will recreate the original tweet(s) in json format, provided the content is still available on Twitter.</p>
#elxn42 tweets (42nd Canadian Federal Election)
<p>Tweet ids for #elxn42 tweets. Tweets can be "hydrated" with Ed Summers' twarc (https://github.com/edsu/twarc). twarc.py --hydrate elxn42-tweet-ids.txt > elxn42-tweets.json. Hydrating will recreate the original tweet(s) in json format, provided the content is still available on Twitter. This dataset is the combination of hydrated http://hdl.handle.net/10864/11310 tweet ids, and htttp://hdl.handle.net/10864/11270.</p>
#MakeDonaldDrumpfAgain tweets
<p>Derivative data for #MakeDonaldDrumpfAgain tweets. Tweets can be "hydrated" with Ed Summers' twarc (https://github.com/edsu/twarc). twarc.py --hydrate MakeDonaldDrumpfAgain-tweet-ids.txt > MakeDonaldDrumpfAgain.json. Hydrating will recreate the original tweet(s) in json format, provided the content is still available on Twitter. This dataset is the combination of hydrated http://hdl.handle.net/10864/11310 tweet ids, and htttp://hdl.handle.net/10864/11270.</p>
Datasets for paper: J. A. Cerón-Guzmán and E. León-Guzmán (2016), A Sentiment Analysis System of Spanish Tweets and Its Application in Colombia 2014 Presidential Election. SocialCom 2016. DOI: 10.1109/BDCloud-SocialCom-SustainCom.2016.47
<p>Datasets for paper: J. A. Cerón-Guzmán and E. León-Guzmán (2016), A Sentiment Analysis System of Spanish Tweets and Its Application in Colombia 2014 Presidential Election. The 9th IEEE International Conference on Social Computing and Networking (SocialCom). DOI: 10.1109/BDCloud-SocialCom-SustainCom.2016.47</p>
#CdnPoli Tweet IDs for 29 December 2016
<p>These are the #CdnPoli TweetIDs for 29 December 2016, collected on 29 December 2016. They were collected using the DocNow project, available at http://docnow.io/. </p>
SAA2017 TAGS Tweet Archive
<p>An open archive of Tweets from SAA2017, the Society for American Archaeology's 82nd Annual Meeting, Vancouver, BC, Canada.</p>
Geo-tagged Tweets in Paris during Nov 2015
<p><strong>Abstract</strong></p> <p>The data sets released here have been used in our study on quantitatively evaluating the impact of disasters in the city. The study of disaster events and their impact in the urban space has been traditionally conducted through manual collections and analysis of surveys, questionnaires and authority documents. While there have been increasingly rich troves of human behavioral data related to the events of interest, the ability to obtain hindsight following a disaster event has not been scaled up. In this study, we propose a novel approach for analyzing events called PairFac. PairFac utilizes discriminant tensor analysis to automatically discover the impact of a major event from rich human behavioral data. Our method aims to (i) uncover the persistent patterns across multiple interrelated aspects of urban behavior (e.g., when, where and what citizens do in a city) and at the same time (ii) identify the salient changes following a potentially impactful event. We show the effectiveness of PairFac in comparison with previous methods through extensive experiments. We also demonstrate the advantages of our approach through case studies with real-world traffic sensor data and social media streams surrounding the 2015 terrorist attacks in Paris. Our work has both methodological contributions in studying the impact of an external stimulus on a system as well as practical implications in the area of disaster event analysis and assessment.</p> <p><strong>Dataset</strong></p> <p>There are two datasets used in this study, traffic sensor dataset and social media dataset. </p> <p>Traffic Sensor dataset was collected from open data Paris. This dataset includes all the hourly data for the flow and the occupancy rate assembled by the permanent traffic sensors installed on the Paris City network (urban network and peripheral boulevard). For interested readers, please direct to the link as: https://opendata.paris.fr/explore/dataset/comptages-routiers-permanents/</p> <p>Social media dataset was collected from Twitter API. The dataset contains geo-tagged tweets from Paris collected through Twitter API between the period of Oct 16th, 2015 and Nov 20, 2015. 75,982 geo-located tweets were extracted during the period covered.</p> <p>Duration: 2015-10-16 to 2015-11-20.</p> <p>Total number of tweets: 75,982</p> <p><strong>Publication</strong></p> <p>If you make use of this data set, please kindly cite:</p> <p>Xidao Wen, Yu-Ru Lin, and Konstantinos Pelechrinis. 2016. PairFac: Event Analytics through Discriminant Tensor Factorization. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management (CIKM '16). ACM, New York, NY, USA, 519-528. DOI: https://doi.org/10.1145/2983323.2983837</p> <p> </p>
TweetC19SR-Eng - Manually annotated dataset of English language COVID-19 tweets containing self-reports of symptoms
<p><strong>In this work, we release two expert curated, manually annotated datasets of COVID-19 self-reported symptoms. The first dataset contains tweets in English and the second contains tweets in Spanish, both containing around 36,500 tweets in total. These datasets were used for the Sixth and Seventh Workshop on Social Media Mining For Health (2021 and 2022)</strong></p>
TweetC19SR-Spa - Manually annotated dataset of Spanish language COVID-19 tweets containing self-reports of symptoms
<p><strong>In this work, we release two expert curated, manually annotated datasets of COVID-19 self-reported symptoms. The first dataset contains tweets in English and the second contains tweets in Spanish, both containing around 36,500 tweets in total. These datasets were used for the Sixth and Seventh Workshop on Social Media Mining For Health (2021 and 2022)</strong></p>
Sample tweets from May 2022
<p>Dataset containing twitter data, namely hashed twitter id, hashed user id, tweet language, user statistics </p><p> </p>
Twitter Dataset - Over 200,000 Tweets containing the word "Vaccine" for research porpuses
<p>This dataset contains 220,085 tweets containing the word vaccine between December 9th and December 18th 2021 at different times during each day, extracted using the Twitter API v2. Each tweet was extracted at least 3 days after its initial posting time in order to register 3 days of engagements, and it doesn't include retweets.</p> <p>Includes:</p> <ul> <li>Tweet ID</li> <li>Text</li> <li>Author ID</li> <li>Date</li> <li>Like count</li> <li>Retweet count</li> <li>Quote count</li> <li>Reply count</li> <li>User data (Followers, Following, Tweet count, Account creation date, Verified status)</li> </ul> <p>Usernames are hidden for privacy reasons</p>
Multilingual MigrationsKB: A Mulitlingual Knowledge Base of Migration related annotated Tweets
<p><strong>Multilingual MigrationskB (MGKB) </strong>is a mulitlingual extended version of English <a href="https://zenodo.org/record/5206820#.YRqF1nUza0o">MGKB</a>. The tweets geotagged with Geo location from 32 European Countries (<em><strong>Austria, Belgium, Bulgaria, Croatia, Cyprus, Czech, Denmark, Estonia, Finland, France, Germany, Greece, Hungary, Ireland, Italy, Latvia, Lithuania, Luxembourg, Malta, Netherlands, Poland, Portugal, Romania, Slovakia, Slovenia, Spain, Sweden, Iceland, Liechtenstein, Norway, Switzerland, the United Kingdom</strong></em>) are extracted and filtered by 11 languages (<em><strong>English, French, Finnish, German, Greek, Dutch, Hungarian, Italian, Polish, Spain, Swedish</strong></em>). Metadata information about the tweets, such as <strong>Geo information (place name, coordinates, country code)</strong> are included. <strong>MGKB</strong> contains <strong>sentiments, offensive and hate speeches, topics, hashtags, user mentions</strong> in RDF format. The schema of <strong>MGKB</strong> is an extension of TweetsKB for migration related information. Moreover, to associate and represent the potential economic and social factors driving the migration flows, the data from <a href="https://ec.europa.eu/eurostat/web/main/home">Eurostat</a> and <a href="https://spec.edmcouncil.org/fibo/ontology/">FIBO</a> ontology was used. To represent multilinguality, the<a href="https://www.cidoc-crm.org/"> CIDOC Conceptual Reference Model (CIDOC-CRM)</a> is used. The extracted economic indicators, i.e., GDP Growth Rate, Total Unemployment Rate, Youth Unemployment Rate, Long-term Unemployment Rate and Income per househould, are connected with each tweet in RDF using geographical and temporal dimensions. </p> <p>For this version, the Multilingual MGKB is delivered separated by year. The extracted topic words are also published.</p> <p>Code: <a href="https://github.com/migrationsKB/MRL">https://github.com/migrationsKB/MRL</a></p> <p>Please contact Yiyi Chen (yiyi.chen@partner.kit.edu) for pretrained models (Sentiment analysis/hate speech detection/ETM) if necessary.</p> <p> </p> <p> </p>
Preprocessed Tweets for LDA with defined Topics
<p>This file contains preprocessed tweets from the #BTW17 Twitter Dataset (Kratzke, 2017). With this tweets, the <code>btw17_twitter_lda.ipynb</code> file from https://github.com/matteomeier/topics-tweets-querysuggestions-btw17 can be run to recreate the results of the analysis for the Bachelor Thesis <strong>Topic-Analyse politischer Tweets und Suchvorschläge zur Bundestagswahl 2017.</strong></p>
MigrationsKB: A Knowledge Base of Migration related annotated Tweets
<p><strong>MigrationsKB(MGKB)</strong> is a public Knowledge Base of anonymized <strong>Migration</strong> related <strong>annotated</strong> tweets. The MGKB currently contains over <strong>200 thousand</strong> tweets, spanning over 9 years (January 2013 to July 2021), filtered with 11 European countries of <em>the United Kingdom, Germany, Spain, Poland, France, Sweden, Austria, Hungary, Switzerland, Netherlands and Italy</em>. <strong>Metadata</strong> information about the tweets, such as Geo information (<strong>place name</strong>, <strong>coordinates</strong>, <strong>country code</strong>). <strong>MGKB</strong> contains <strong>entities</strong>, <strong>sentiments</strong>, <strong>hate speeches</strong>, <strong>topics</strong>, <strong>hashtags</strong>, <em>encrypted user mentions</em> in RDF format. The schema of <strong>MGKB</strong> is an extension of TweetsKB for migrations related information. Moreover, to associate and represent the potential economic and social factors driving the migration flows such as <a href="https://ec.europa.eu/eurostat/web/main/home"><strong>eurostat</strong></a>, <a href="https://www.statista.com/"><strong>statista</strong></a>, etc. FIBO ontology was used. The extracted <strong>economic indicators</strong>, such as GDP Growth Rate, are connected with each Tweet in RDF using geographical and temporal dimensions. The user IDs and the tweet texts are encrypted for privacy purposes, while the tweet IDs are preserved.</p> <p>For this version, the <strong>MGKB</strong> is delivered as a whole and separately by year. The extracted entities and topic words are also published.</p> <p>Online SPARQL endpoint <a href="https://mgkb.fiz-karlsruhe.de/sparql/">https://mgkb.fiz-karlsruhe.de/sparql/</a></p> <p>More information please refer to the website <a href="https://migrationskb.github.io/MGKB/">https://migrationskb.github.io/MGKB/</a>.</p> <p>Please contact Yiyi Chen (yiyi.chen@partner.kit.edu) for pretrained models (sentiment analysis/hate speech detection/ETM) if necessary.</p>
Twitter Timelines of British MPs (Tweet IDs)
<p>TSV file containing Twitter tweet_ids from the Timelines of 584 members of British Parliament (collected between 4th and 6th of March 2022). The users were identified from the link below:</p> <p>https://www.ukinbound.org/resources/list-of-mp-twitter-accounts/</p> <p> </p> <p>If you use this dataset for an academic work, please reference the following paper:</p> <pre>@article{tacchi2022signed, title={Signed ego network model and its application to Twitter}, author={Tacchi, Jack and Boldrini, Chiara and Passarella, Andrea and Conti, Marco}, journal={arXiv preprint arXiv:2206.15228}, year={2022} }</pre>
Brazilian tweets classified for sentiment analysis
<p>Brazilian tweets classified for sentiment analysis</p>
SocialDisNER corpus: gold standard annotations for detection of disease mentions in Spanish tweets
<p><strong>If you use any data from this repository, please cite our scientific paper instead of the Zenodo repo: </strong></p> <p>Luis Gasco Sánchez, Darryl Estrada Zavala, Eulàlia Farré-Maduell, Salvador Lima-López, Antonio Miranda-Escalada, and Martin Krallinger. 2022. <a href="https://aclanthology.org/2022.smm4h-1.48">The SocialDisNER shared task on detection of disease mentions in health-relevant content from social media: methods, evaluation, guidelines and corpora</a>. In <em>Proceedings of The Seventh Workshop on Social Media Mining for Health Applications, Workshop & Shared Task</em>, pages 182–189, Gyeongju, Republic of Korea. Association for Computational Linguistics.</p> <pre><code class="language-json">@inproceedings{gasco2022socialdisner, title = "The {S}ocial{D}is{NER} shared task on detection of disease mentions in health-relevant content from social media: methods, evaluation, guidelines and corpora", author = "Gasco S{\'a}nchez, Luis and Estrada Zavala, Darryl and Farr{\'e}-Maduell, Eul{\`a}lia and Lima-L{\'o}pez, Salvador and Miranda-Escalada, Antonio and Krallinger, Martin", booktitle = "Proceedings of The Seventh Workshop on Social Media Mining for Health Applications, Workshop {\&} Shared Task", month = oct, year = "2022", address = "Gyeongju, Republic of Korea", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2022.smm4h-1.48", pages = "182--189" }</code></pre> <p> </p> <p><strong>Introduction:</strong><br> The <strong>SocialDisNER corpus</strong> of the SMM4H 2022 – Task 10 task focus on the recognition of disease mentions in tweets written in Spanish after selecting primarily<strong><em> first-hand experience of diseases</em></strong> and other health-relevant content (from patient associations, professional healthcare institutions, and through followers of patient association accounts of a <em>diversity of pathologies</em> including rare diseases, mental health, cancer, etc..).</p> <p><strong>SocialDisNER Gold Standard</strong></p> <p>The Gold Standard corpus was manually annotated by medical experts following the <a href="https://doi.org/10.5281/zenodo.6983041">SMM4H-SocialDisNER guidelines</a>. These guidelines were adapted from previous efforts used to annotate patient clinical records and medical literature. It covers rules for annotating <strong>mentions of diseases</strong> in health-related tweets in Spanish,</p> <p>The training set consists of 5000 tweets written in Spanish and the validation set consists of 2500 tweets written in Spanish. Both sets have been manually annotated by healthcare professionals. The test dataset contains 23430 tweets, although only 2000 will be used to evaluate the systems participating in the task (the rest is background set). We don't plan to publish the test set, but if you want you can test your system from <a href="https://codalab.lisn.upsaclay.fr/competitions/3531">SocialDisNER Codalab</a>.</p> <p><strong>SocialDisNER Large Scale Corpus</strong></p> <p>The large-scale data contains mentions automatically extracted from a set of 85000 tweets. Separate datasets are shown for each entity including diseases, drugs, symptoms, professions, procedures, species, morphology neoplasm, and persons.</p> <p><strong>SocialDisNER co-mention networks</strong></p> <p>We have computed a co-occurrence matrix of the extracted diseases, as well as several co-mention matrices between the disease mentions and the rest of the entities in the large-scale corpora.</p> <p> </p> <p><strong>File structure:</strong></p> <p>The structure of the corpus is: </p> <ul> <li><strong>SocialDisNER_Data:</strong> <ul> <li>training-validation-data folder <ul> <li><strong><em>train-valid-txt-files</em></strong>: folder with training and validation text files. One text file per tweet, the file name corresponds to the tweet id. One sub-directory per corpus split (train and valid). The files named <em>ids_dev_set.txt</em> and<em> ids_train_set.txt </em>contain the list of file identifiers for each of the data splits (validation and train).</li> <li><strong><em>mentions.tsv</em></strong>: This file contains the manually annotated disease mentions. The file has the following fields: <ul> <li><em>tweets_id</em>: This is the id of the tweet, using Twitter API you can query the content of the tweet.</li> <li><em>Begin</em>: This is the position in the tweet where the annotation was found.</li> <li><em>End</em>: This is the position of the last character of the annotation in the tweet.</li> <li><em>Type: </em>This is the type of entity found, in our case "ENFERMEDAD".</li> <li><em>Extraction</em>: This is the literal extraction, in other words, the fragment of text which refers to the annotation. </li> </ul> </li> </ul> </li> <li>test-data folder: <ul> <li><strong>test-data-txt-files</strong>: folder with test text files. One file per tweet, the file name corresponds to the tweet id. The folder contains 23430 tweets to be used as test set of the task. Of them, 2000 will be used to evaluate the participating systems.</li> </ul> </li> </ul> </li> </ul> <p> </p> <ul> <li><strong>SocialDisNER_LargeScale_additionaldata:</strong> <ul> <li>socialdisner_diseases: <ul> <li><strong>tweets_txt:</strong> Folder with large-scale tweet database. One text file per tweet, the file name corresponds to the tweet id.</li> <li><strong>diseases_mentions.tsv</strong>: This file contains the automatically annotated disease mentions from the large-scale SocialDisNER corpus (Silver Standard). The structure is the same than the Golden Standard annotations.</li> </ul> </li> <li>socialdisner_ENTITY: Each folder with this naming convention contains the following data structure. Corpora have been generated with mentions of diseases, drugs, symptoms, professions, procedures, species, morphology neoplasm and persons <ul> <li><strong>tweets_txt:</strong> Folder with large-scale tweet database. One text file per tweet, the file name corresponds to the tweet id.</li> <li><strong>ENTITY_mentions.tsv</strong>: This file contains the automatically annotated mentions of type “ENTITY” from the large-scale SocialDisNER corpus (Silver Standard). The structure is the same than the Golden Standard annotations.</li> </ul> </li> <li>socialdisner_networks: This folder contains tsv files containing the co-mention matrices between the diseases and the rest of the entities of the large-scale socialdisner data. Each file follows the following naming convention: <ul> <li><strong>socialdisner_disease-ENTITY_net.tsv</strong><em>: </em>The tsv file contains a series of columns and rows corresponding to the mentions used for building the matrix. Each column is separated by “;”. The type of each mention is identified by the label in parentheses of each title. The count represents the number of times that mention x and mention y were found in the same tweet of the large-scale dataset.</li> <li><strong>socialdiser_disease_net.tsv</strong>: This tsv file contains the array of socialdisner-disease large-scale corpus co-mentions separated by ";". This file can be loaded into NetworkX to perform disease co-morbidity analysis on the socialdisner-disease large-scale data.</li> </ul> </li> </ul> </li> </ul> <p><em>Note: In previous versions of the dataset the order of the columns in the mentions.tsv file was not in the correct order. From this version onwards the order is correct and adequate to send the predictions of the task.</em></p> <p> </p> <p>For further information, please visit <a href="https://temu.bsc.es/socialdisner/">https://temu.bsc.es/socialdisner/</a></p> <p><strong>Summary statistics:</strong></p> <table> <caption>Manually annotated data</caption> <thead> <tr> <th scope="row"> </th> <th scope="col">Training set</th> <th scope="col">Development set</th> </tr> </thead> <tbody> <tr> <th scope="row"># tweets</th> <td>5000</td> <td>2500</td> </tr> <tr> <th scope="row"># characters</th> <td>1253431</td> <td>516768</td> </tr> <tr> <th scope="row"># tokens</th> <td>211555</td> <td>84478</td> </tr> <tr> <th scope="row">Avg. char / tweet</th> <td>250.69</td> <td>206.71</td> </tr> <tr> <th scope="row">Avg. tok. / tweet</th> <td>42.31</td> <td>33.79</td> </tr> <tr> <th scope="row"># mentions</th> <td>15173</td> <td>4252</td> </tr> <tr> <th scope="row"># unique mentions</th> <td>4407</td> <td>1413</td> </tr> </tbody> </table> <p> </p> <table> <caption>Large-scale annotated data (Silver Standard)</caption> <tbody> <tr> <td> </td> <td><em>Socialdisner-diseases</em></td> <td><em>Socialdisner-pharma</em></td> <td><em>Socialdisner-morphology_neoplasms</em></td> <td><em>Socialdisner-symptoms</em></td> <td><em>Socialdisner-professions</em></td> <td><em>Socialdisner-Procedures</em></td> <td><em>Socialdisnerv-Person</em></td> <td><em>Socialdisner-Species</em></td> </tr> <tr> <td><strong># tweets</strong></td> <td>85077</td> <td>1759</td> <td>8518</td> <td>12624</td> <td>15831</td> <td>11462</td> <td>41033</td> <td>12118</td> </tr> <tr> <td><strong># characters</strong></td> <td>19920670</td> <td>435141</td> <td>2082574</td> <td>3023784</td> <td>4063114</td> <td>2873791</td> <td>10273278</td> <td>2933925</td> </tr> <tr> <td><strong># tokens</strong></td> <td>3236411</td> <td>68269</td> <td>332539</td> <td>521503</td> <td>660071</td> <td>467059</td> <td>1689479</td> <td>486249</td> </tr> <tr> <td><strong>Avg. char / tweet</strong></td> <td>234.15</td> <td>247.38</td> <td>244.49</td> <td>239.53</td> <td>256.66</td> <td>250.72</td> <td>250.37</td> <td>242.11</td> </tr> <tr> <td><strong>Avg. tok. / tweet</strong></td> <td>38.04</td> <td>38.81</td> <td>39.04</td> <td>41.31</td> <td>41.69</td> <td>40.75</td> <td>41.17</td> <td>40.13</td> </tr> <tr> <td><strong># mentions</strong></td> <td>116260</td> <td>1029</td> <td>8943</td> <td>12896</td> <td>18590</td> <td>10080</td> <td>58007</td> <td>14014</td> </tr> <tr> <td><strong># unique mentions</strong></td> <td>16034</td> <td>530</td> <td>541</td> <td>6991</td> <td>3667</td> <td>3841</td> <td>3446</td> <td>1676</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p>Do not share the data with other individuals/teams without permission from the task organizer. Tweets IDs are the primary source of information. Tweet texts are provided as support material. By downloading this resource, you agree to the Twitter <a href="https://twitter.com/en/tos">Terms of Service</a>, <a href="https://twitter.com/en/privacy">Privacy Policy</a>, <a href="https://developer.twitter.com/en/developer-terms/agreement">Developer Agreement</a>, and <a href="https://developer.twitter.com/en/developer-terms/policy">Developer Policy</a>.</p> <p> </p> <p> </p>
Labeled data and models for COVID-19 vaccine related tweets with stance, location, and topics
<p>The dataset contains Tweet IDs along with the location and tweet timestamp. The tweets are labeled based on motivating/demotivating status, stance towards the COVID-19 vaccine, and topic in the tweet text. To comply with Twitter guidelines, we removed the tweet texts and author information. You can use Hydrator API to hydrate the tweets.</p> <p>The repository also contains the machine-learning models for topic modeling, de/motivation classifier, and stance detection from the tweets.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.