Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “rumors”
FaCov Dataset: COVID-19 Viral News and Rumors Fact-Check Articles Dataset
<p>The data were collected by web-scraping pages from the websites collected earlier, using the <a href="https://webscraper.io/">Web Scraper browser extension</a>.</p> <p>More specifically, the sections of these websites that dealt exclusively with COVID-19 related content were scraped. In cases where the website did not have such a specified section, the search functionality within the website was used to query terms related to COVID-19 and the articles in the search results were scraped. Also in some cases, all articles were scraped and those unrelated to COVID-19 were filtered out in the pre-processing stage. All the samples collected were then put together into one CSV</p> <p>The following information was extracted along with the articles:</p> <p>Title of the fact check article</p> <p>URL of the fact check article</p> <p>Claim being discussed in the article (if available)</p> <p>Summary of the fact check article (if available)</p> <p>Content of the fact check article• Label assigned by the article to the claim</p> <p>Author of the fact check article (if available)</p> <p>Date of publication of the article (if available)</p>
Linguistic features of Twitter rumor propagation trees on the CLNews19-20 dataset
<p>Linguistic features applied to 140 Twitter rumor propagation trees (53 of type 1: false rumors, and 87 of type 2: true rumors) about Chilean topics collected during the Chilean social outbreak (2019-2020). These rumor propagation trees come from the CLNews19-20 dataset (DOI <a href="../doi/10.5281/zenodo.5851204">10.5281/zenodo.5851204</a>). There are 38 different linguistics features applied:</p> <ul> <li>number of paragraphs</li> <li>total number of sentences</li> <li>standard deviation of sentences</li> <li>mean words</li> <li>maximum number of words</li> <li>mean characters</li> <li>maximum number of characters</li> <li>mean characters without spaces</li> <li>maximum number of characters without spaces</li> <li>adjective idf</li> <li>minimum number of adpositions</li> <li>maximum number of adpositions</li> <li>mean adpositions</li> <li>median adpositions</li> <li>maximum number of auxiliaries</li> <li>total number of auxiliaries</li> <li>mean auxiliaries</li> <li>median auxiliaries</li> <li>auxiliary idf</li> <li>auxiliary tfidf</li> <li>standard deviation of numerals</li> <li>proper noun idf</li> <li>minimum number of symbols</li> <li>maximum number of symbols</li> <li>total number of symbols</li> <li>mean symbols</li> <li>median symbols</li> <li>symbol idf</li> <li>symbol tfidf</li> <li>number of paragraphs</li> <li>MDT conditionals</li> <li>MDT counterarguments</li> <li>MDT Connectors Opinion Justifiers</li> <li>MDT Connectors Opinion Generalizers</li> <li>AS VeryPositive Affin</li> <li>AS Negative Nrc</li> <li>AS Angry Nrc</li> <li>AS Fear Nrc</li> </ul>
Newly Emerged Rumors in Twitter
<p><strong>*** Newly Emerged Rumors in Twitter</strong> <strong>***</strong></p> <p>These 12 datasets are the results of an empirical study on the spreading process of newly emerged rumors in Twitter. Newly emerged rumors are those rumors whose rise and fall happen in a short period of time, in contrast to the long standing rumors. Particularly, we have focused on those newly emerged rumors which have given rise to an anti-rumor spreading simultaneously against them. The story of each rumor is as follow :</p> <p>1- Dataset_R1 : The National Football League team in Washington D.C. changed its name to Redhawks.</p> <p>2- Dataset_R2 : A Muslim waitress refused to seat a church group at a restaurant, claiming "religious freedom" allowed her to do so.</p> <p>3- Dataset_R3 : Facebook CEO Mark Zuckerberg bought a "super-yacht" for $150 million.</p> <p>4- Dataset_R4 : Actor Denzel Washington said electing President Trump saved the U.S. from becoming an "Orwellian police state."</p> <p>5- Dataset_R5 : Joy Behar of "The View" sent a crass tweet about a fatal fire in Trump Tower.</p> <p>6- Dataset_R6 : Harley-Davidson's chief executive officer Matthew Levatich called President Trump "a moron."</p> <p>7- Dataset_R7 : The animated children's program 'VeggieTales' introduced a cannabis character in August 2018.</p> <p>8- Dataset_R8 : Michael Jordan resigned from the board at Nike and took his Air Jordan line of apparel with him.</p> <p>9- Dataset_R9 : In September 2018, the University of Alabama football program ended its uniform contract with Nike, in response to Nike's endorsement deal with Colin Kaepernick.</p> <p>10- Dataset_R10 : During confirmation hearings for Supreme Court nominee Brett Kavanaugh, congressional Democrats demanded that the nominee undergo DNA testing to prove he is not Adolf Hitler.</p> <p>11- Dataset_R11 : Singer Michael Bublé's upcoming album will be his last, as he is retiring from making music.Singer Michael Bublé's upcoming album will be his last, as he is retiring from making music.</p> <p>12- Dataset_R12 : A screenshot from MyLife.com confirms that mail bomb suspect Cesar Sayoc was registered as a Democrat.</p> <p> </p> <p>The structure of excel files for each dataset is as follow :</p> <p>- Each row belongs to one captured tweet/retweet related to the rumor, and each column of the dataset presents a specific information about the tweet/retweet. These columns from left to right present the following information about the tweet/retweet : </p> <p>- User ID (user who has posted the current tweet/retweet)</p> <p>- The description sentence in the profile of the user who has published the tweet/retweet</p> <p>- The number of published tweet/retweet by the user at the time of posting the current tweet/retweet</p> <p>- Date and time of creation of the the account by which the current tweet/retweet has been posted </p> <p>- Language of the tweet/retweet</p> <p>- Number of followers </p> <p>- Number of followings (friends)</p> <p>- Date and time of posting the current tweet/retweet</p> <p>- Number of like (favorite) the current tweet had been acquired before crawling it</p> <p>- Number of times the current tweet had been retweeted before crawling it</p> <p>- Is there any other tweet inside of the current tweet/retweet (for example this happens when the current tweet is a quote or reply or retweet)</p> <p>- The source (OS) of device by which the current tweet/retweet was posted</p> <p>- Tweet/Retweet ID</p> <p>- Retweet ID (if the post is a retweet then this feature gives the ID of the tweet that is retweeted by the current post)</p> <p>- Quote ID (if the post is a quote then this feature gives the ID of the tweet that is quoted by the current post)</p> <p>- Reply ID (if the post is a reply then this feature gives the ID of the tweet that is replied by the current post)</p> <p>- Frequency of tweet occurrences which means the number of times the current tweet is repeated in the dataset (for example the number of times that a tweet exists in the dataset in the form of retweet posted by others)</p> <p>- State of the tweet which can be one of the following forms (achieved by an agreement between the annotators) :</p> <p> r : The tweet/retweet is a rumor post</p> <p> a : The tweet/retweet is an anti-rumor post</p> <p> q : The tweet/retweet is a question about the rumor, however neither confirm nor deny it</p> <p> n : The tweet/retweet is not related to the rumor (even though it contains the queries related to the rumor, but does not refer to the rumor)</p> <p> </p> <p> </p> <p> </p> <p> </p>
CLNews19-20: A new dataset for rumor detection in Spanish
<p>We create CLNews, a dataset for rumor detection in Spanish. Based on fact-checking agencies' data, we mapped related tweets to verify news into four categories: non-rumor, true rumor, false rumor, and unverified rumors. Mapping these news to Twitter, we collected data, including tweet timelines and conversational threads. We release this dataset to promote research in this area in Spanish.</p>
Twitter Data on some Rumors during 2014 Flood Tragedy in Malaysia
<p>Data was collected from Twitter Advanced Search webpage using keyword based on some rumors during period of December 2014 until January 2015</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.