Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

31

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

31 results for “Fake News”

Learn how ShareScore rates datasets ↗
zenodo48/100

BuzzFeed-Webis Fake News Corpus 2016

<p>The corpus comprises the output of 9 publishers in a week close to the US elections. Among the selected publishers are 6 prolific hyperpartisan ones (three left-wing and three right-wing), and three mainstream publishers (see Table 1). All publishers earned Facebook&rsquo;s blue checkmark, indicating authenticity and an elevated status within the network. For seven weekdays (September 19 to 23 and September 26 and 27), every post and linked news article of the 9 publishers was fact-checked by professional journalists at BuzzFeed. In total, 1,627 articles were checked, 826 mainstream, 256 left-wing and 545 right-wing. The imbalance between categories results from differing publication frequencies.</p>

opencc-by-4.0Feb 2018View details →
zenodo44/100

FA-KES: A Fake News Dataset around the Syrian War

<p>We have produced a labeled dataset that presents fake news surrounding the conflict in Syria. The dataset consists of a set of articles/news labeled by 0 (fake) or 1 (credible). Credibility of articles are computed with respect to a ground truth information obtained from the Syrian Violations Documentation Center&nbsp; (VDC). In particular, for each article, we crowdsource the information extraction (e.g., date, location, Number of casualties) job using the crowdsourcing platform Figure Eight (formally CrowdFlower). Then, we match those articles against the VDC database to be able to deduce whether an article is fake or not. The dataset can be used to train machine learning models to detect fake news.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

South African Disinformation [Fake News] Website Data - 2020

<p>See publication:&nbsp;<strong>Is it Fake? News Disinformation Detection on South African News Websites</strong></p> <p>We used, as sources, investigations by the news websites MyBroadband (<a href="https://mybroadband.co.za/forum/threads/list-of-known-fake-news-sites-in-south-africa-and-beyond.879854/">https://mybroadband.co.za/forum/threads/list-of-known-fake-news-sites-in-south-africa-and-beyond.879854/</a>) and News24 (<a href="https://exposed.news24.com/the-website-blacklist/">https://exposed.news24.com/the-website-blacklist/</a>). These articles covered investigations into disinformation websites in South Africa in 2018. They compiled lists of websites that were suspected to be disinformation. During the period from those articles to present, a number of the websites have become inaccessible or offline. We attempted to use the internet archives <a href="https://archive.org/web/">WayBack Machine</a>&nbsp;we could only get partial snapshots and error messages.</p> <p>A web-scraper only worked for one of the sources although manual editing was still required to clean the text from Javascript code and some paragraph duplicates. On most of the other websites, a web-scraper did not work well as there were too many advertisements and broken parts of pages. Because of all these problems, most of the articles were manually copied and pasted and cleaned in flat files. In some cases, the text of articles could not be copied and was not made part of the South African disinformation corpus.</p> <p><strong>Citing the dataset</strong></p> <blockquote> <p>@inproceedings{de2021fake, title={Is it Fake? News Disinformation Detection on South African News Websites}, author={de Wet, Harm and Marivate, Vukosi}, booktitle={2021 IEEE AFRICON}, pages={1--6}, year={2021}, organization={IEEE} }</p> </blockquote>

opencc-by-sa-4.0Jul 2021View details →
zenodo40/100

COVID Fake News Dataset

<p><strong>Context</strong></p> <p>The dataset contains the list of COVID Fake News/Claims which is shared all over the internet.</p> <p><strong>Content</strong></p> <ol> <li>Headlines: String attribute consisting of the headlines/fact shared.</li> <li>Outcome: It is binary data where 0 means the headline is fake and 1 means that it is true.</li> </ol> <p><strong>Inspiration</strong></p> <p>In many research portals, there was this common question in which the combined fake news dataset is available or not. This led to the publication of this dataset.</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Mitigating Biases in Collective Decision-Making: Enhancing Performance in the Face of Fake News

<p>Data supporting "Mitigating Biases in Collective Decision-Making: Enhancing Performance in the Face of Fake News".&nbsp;<br><br></p> <p>&nbsp;If you use this dataset in your own research, please cite this paper:</p> <p>```<br>@misc{abels2024mitigating,<br>&nbsp; &nbsp; &nbsp; title={Mitigating Biases in Collective Decision-Making: Enhancing Performance in the Face of Fake News},&nbsp;<br>&nbsp; &nbsp; &nbsp; author={Axel Abels and Elias Fernandez Domingos and Ann Now&eacute; and Tom Lenaerts},<br>&nbsp; &nbsp; &nbsp; year={2024},<br>&nbsp; &nbsp; &nbsp; eprint={2403.08829},<br>&nbsp; &nbsp; &nbsp; archivePrefix={arXiv},<br>&nbsp; &nbsp; &nbsp; primaryClass={cs.HC}<br>}<br>```</p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>column name</strong></td> <td><strong>description</strong></td> </tr> <tr> <td>treatment</td> <td>identifier for the set of headlines presented to the participant</td> </tr> <tr> <td>trial</td> <td>trial/round in which the headline was presented&nbsp;</td> </tr> <tr> <td>arm</td> <td>which "arm" the headline was presented as (0=left, 1=middle, 2=right)</td> </tr> <tr> <td>advice</td> <td>the participant's response (0=very unlikely, 0.25=unlikely, 0.5=undecided, 0.75=likely, 1=very likely)</td> </tr> <tr> <td>genuine</td> <td>whether the headline was genuine (1) or altered (0)</td> </tr> <tr> <td>headline</td> <td>the headline as shown to the participant</td> </tr> <tr> <td>original</td> <td>the headline before a possible alteration</td> </tr> <tr> <td>expert_id</td> <td>participant's identifier</td> </tr> <tr> <td>sentiment</td> <td>whether the headline reported a negative (-1) or positive (1) outcome</td> </tr> <tr> <td>expert:ethnicity</td> <td>the participant's ethnicity</td> </tr> <tr> <td>expert:sex</td> <td>the participant's sex</td> </tr> <tr> <td>expert:age</td> <td>the participant's age</td> </tr> <tr> <td>outcome:white, outcome:black, outcome:young, outcome:old, outcome:male, outcome:female</td> <td>whether the headline reported a negative (-1) or positive (1) or neutral (0) outcome for the specified group</td> </tr> <tr> <td>trial_time</td> <td>how long the participant took to respond to the trial/round</td> </tr> </tbody> </table> <p><strong>abstract</strong><br>Individual and social biases undermine the effectiveness of human advisers by inducing judgment errors which can disadvantage protected groups. In this paper, we study the influence these biases can have in the pervasive problem of fake news by evaluating human participants' capacity to identify false headlines. By focusing on headlines involving sensitive characteristics, we gather a comprehensive dataset to explore how human responses are shaped by their biases. Our analysis reveals recurring individual biases and their permeation into collective decisions. We show that demographic factors, headline categories, and the manner in which information is presented significantly influence errors in human judgment. We then use our collected data as a benchmark problem on which we evaluate the efficacy of adaptive aggregation algorithms. In addition to their improved accuracy, our results highlight the interactions between the emergence of collective intelligence and the mitigation of participant biases.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Propaganda and fake news on the war in Ukraine: data from Russian-speaking social media communities

<p>The data set contains posts from social media networks popular among Russian-speaking communities. Information was searched based on pre-defined keywords (&quot;war&quot;, &quot;special military operation&quot;,&nbsp;etc.) and is mainly related to the ongoing war in Ukraine with Russia. After a thorough review and analysis of the data, both propaganda and fake news were identified.&nbsp;The collected data is anonymized. Feature engineering and text preprocessing can be applied to obtain new insights and knowledge from this data set. The data set is useful for the study of information wars and propaganda identification.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Fake News

<p>This data collection focuses on capturing user-generated content from the popular social network Reddit in 2024. The dataset &ldquo;Fake News&rdquo; comprises collected data from 3636 users of Reddit. This dataset consists of .csv .xls, and .xlsx files, containing textual data associated with fake news.</p> <p>Funded by the EU NextGeneration EU through the Recovery and Resilience Plan for Slovakia under the project No. 09I03-03-V01-000153</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Datasets from Costa Rican news sources for fake news detection

<p>Today, technology has changed the way information is propagated and how the message is received. The interpretation of the news may have different angles depending on the source of origin. Because of this, there has been an increase in misinformation, in the way of influencing public opinion and in how we perceive or estimate reality.<br> The objective of this beta dataset is to be used for the evaluation of data mining models that allow the classification of true or potentially fake news that are generated by Costa Rican news sites only.&nbsp; This is intended to assess the level of reliability of the models and extend the scope of this research in future work.</p> <p>The dataset has been pre-processed (standarized using lower cases, lemmatized and removed any possible noise from it) and analyzed using LIWC dictionaries. One version has the news text in Spanish and was processed using LIWC2007 dictionary in Spanish. The second version was processed using LIWC2015 English dictionary and has the news text in English. The reason to having two versions is to be able to test using the newer LIWC dictionary which includes more Summary Language Variables that the Spanish version doesn&#39;t have and analyze how this and other variables can contribute to different results when creating models.&nbsp;</p> <p>The file &quot;DescripcionVariables&quot; provides a description of all variables used.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

A study on real graphs of fake news spreading on Twitter

<p><strong>*** Fake News on Twitter</strong>&nbsp;<strong>***</strong></p> <p>These 5 datasets are the results of an empirical study on the spreading process of newly fake news on Twitter. Particularly, we have focused on those fake news which have given rise to a truth spreading simultaneously against them. The story of each fake news is as follow:</p> <p>1-&nbsp;FN1:&nbsp;A Muslim waitress refused to seat a church group at a restaurant, claiming &quot;religious freedom&quot; allowed her to do so.</p> <p>2-&nbsp;FN2:&nbsp;Actor Denzel Washington said electing President Trump saved the U.S. from becoming an &quot;Orwellian police state.&quot;</p> <p>3-&nbsp;FN3:&nbsp;Joy Behar of &quot;The View&quot; sent a crass tweet about a fatal fire in Trump Tower.</p> <p>4-&nbsp;FN4:&nbsp;The animated children&#39;s program &#39;VeggieTales&#39; introduced a cannabis character in August 2018.</p> <p>5-&nbsp;FN5:&nbsp;In September 2018, the University of Alabama football program ended its uniform contract with Nike, in response to Nike&#39;s endorsement deal with Colin Kaepernick.</p> <p>The data collection has been done in two stages that each provided a new dataset: 1- attaining Dataset of Diffusion (DD) that includes information of fake news/truth tweets and retweets 2- Query of neighbors for spreaders of tweets that provides us with Dataset of Graph (DG).&nbsp;</p> <p><strong>DD </strong></p> <p>DD for each fake news story is an excel file, named FNx_DD where x is the number of fake news, and has the following structure:</p> <p>The structure of excel files for each dataset is as follow:</p> <ul> <li>Each row belongs to one captured tweet/retweet related to the rumor, and each column of the dataset presents a specific information about the tweet/retweet. These columns from left to right present the following information about the tweet/retweet:&nbsp;</li> <li>User ID (user who has posted the current tweet/retweet)</li> <li>The number of published tweet/retweet by the user at the time of posting the current tweet/retweet</li> <li>Language of the tweet/retweet</li> <li>Number of followers&nbsp;</li> <li>Number of followings (friends)</li> <li>Date and time of posting the current tweet/retweet</li> <li>Number of like (favorite) the current tweet had been acquired before crawling it</li> <li>Number of times the current tweet had been retweeted before crawling it</li> <li>Is there any other tweet inside of the current tweet/retweet (for example this happens when the current tweet is a quote or reply or retweet)</li> <li>The source (OS) of device by which the current tweet/retweet was posted</li> <li>Tweet/Retweet ID</li> <li>Retweet ID (if the post is a retweet then&nbsp;this feature gives the ID of the tweet that is retweeted by the current post)</li> <li>Quote ID (if the post is a quote then this feature gives the ID of the tweet that is quoted by the current post)</li> <li>Reply ID (if the post is a reply then this feature gives the ID of the tweet that is replied by the current post)</li> <li>Frequency of tweet occurrences which means the number of times the current tweet is repeated in the dataset (for example the number of times that a tweet exists in the dataset in the form of retweet posted by others)</li> <li>State of the tweet which can be one of the following forms (achieved by an agreement between the annotators):</li> </ul> <ul> <li>r : The tweet/retweet is a fake news post</li> <li>a : The tweet/retweet is a truth post</li> <li>q : The tweet/retweet is a question about the fake news, however neither confirm nor deny it</li> <li>n : The tweet/retweet is not related to the fake news (even though it contains the queries related to the rumor, but does&nbsp;&nbsp; not refer&nbsp;to&nbsp;the&nbsp;given fake news)</li> </ul> <p>&nbsp;</p> <p><strong>DG</strong></p> <p>DG for each fake news contains two files:</p> <ul> <li>A file in graph format (.graph) which includes the information of graph such as who is linked to whom. (This file named FNx_DG.graph, where x is the number of fake news)</li> <li>A file in Jsonl format (.jsonl) which includes the real user IDs of nodes in the graph file. (This file named FNx_Labels.jsonl, where x is the number of fake news)</li> </ul> <p>Because in the graph file, the label of each node is the number of its entrance in the graph. For example if node with user ID 12345637 be the first node which has been entered into the graph file then its label in the graph is 0 and its real ID (12345637) would be at the row number 1 (because the row number 0 belongs to column labels) in the jsonl file and so on other node IDs would be at the next rows of the file (each row corresponds to 1 user id). Therefore, if we want to know for example what the user id of node 200 (labeled 200 in the graph) is, then in jsonl file we should look at row number 202.</p> <p>&nbsp;</p> <p>The user IDs of spreaders in DG (those who have had a post in DD) would be available in DD to get extra information about them and their tweet/retweet. The other user IDs in DG are the neighbors of these spreaders and might not exist in DD.</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

Fake News Text Collections

<p>The description are in:&nbsp;<a href="https://github.com/GoloMarcos/FKTC">https://github.com/GoloMarcos/FKTC</a></p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

COVID-19 Fake News Detection Dataset

<p>Please note that this data set was originally shared by Patwa et al. (2021) on GitHub.&nbsp;</p> <p><strong>Reference </strong></p> <p>Patwa, P., Sharma, S. Pykl, S., Guptha, V., Kumari, G., Akhtar, M. S., Ekbal, A., Das A. &amp; Chakraborty, T. (2021). Fighting an Infodemic: COVID-19 Fake News Dataset. Combating Online Hostile Posts in Regional Languages during Emergency Situation, Cham, Springer International Publishing. https://doi.org/10.1007/978-3-030-73696-5_3</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Japanese Fake News Dataset

<p>Japanese Fake News Dataset</p> <p>Read our project page for more details:&nbsp;https://hkefka385.github.io/dataset/fakenews-japanese/</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

CT-FAN: A Multilingual dataset for Fake News Detection

<p><strong>By downloading the data, you agree with the terms &amp; conditions mentioned below:</strong></p> <p><strong>Data Access:&nbsp;</strong>The data in the research collection may only be used for research purposes. Portions of the data are copyrighted and have commercial value as data, so you must be careful to use them only for research purposes.&nbsp;</p> <p>Summaries, analyses and interpretations of the linguistic properties of the information may be derived and published, provided it is impossible to reconstruct the information from these summaries.&nbsp;You may not try identifying the individuals whose texts are included in this dataset. You may not try to identify the original entry on the fact-checking site. You are not permitted to publish any portion of the dataset besides summary statistics or share it with anyone else.</p> <p>We grant you the right to access the collection&#39;s content as described in this agreement. You may not otherwise make unauthorised commercial use of, reproduce, prepare derivative works, distribute copies, perform, or publicly display the collection or parts of it. You are responsible for keeping and storing the data in a way that others cannot access. The data is provided free of charge.</p> <p><strong>Citation</strong></p> <p>Please cite our work&nbsp;as</p> <pre>@InProceedings{clef-checkthat:2022:task3, author = {K{\&quot;o}hler, Juliane and Shahi, Gautam Kishore and Stru{\ss}, Julia Maria and Wiegand, Michael and Siegel, Melanie and Mandl, Thomas}, title = &quot;Overview of the {CLEF}-2022 {CheckThat}! Lab Task 3 on Fake News Detection&quot;, year = {2022}, booktitle = &quot;Working Notes of CLEF 2022---Conference and Labs of the Evaluation Forum&quot;, series = {CLEF~&#39;2022}, address = {Bologna, Italy},} </pre> <pre>@article{shahi2021overview, title={Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection}, author={Shahi, Gautam Kishore and Stru{\ss}, Julia Maria and Mandl, Thomas}, journal={Working Notes of CLEF}, year={2021} }</pre> <p><strong>Problem Definition:</strong> Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other (e.g., claims in dispute) and detect the topical domain of the article. This task will run in <strong>English and German.</strong></p> <p><strong>Task 3:</strong> <strong>Multi-class fake news detection of news articles (English)</strong>&nbsp;Sub-task A would detect fake news designed as a four-class classification problem. Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other. The training data will be released in batches and roughly about 1264 articles with the respective label in English language. Our definitions for the categories are as follows:</p> <ul> <li> <p>False - The main claim made in an article is untrue.</p> </li> <li> <p>Partially False - The main claim of an article is a mixture of true and false information. The article contains partially true and partially false information but cannot be considered 100% true. It includes all articles in categories like partially false, partially true, mostly true, miscaptioned, misleading etc., as defined by different fact-checking services.</p> </li> <li> <p>True - This rating indicates that the primary elements of the main claim are demonstrably true.</p> </li> <li> <p>Other- An article that cannot be categorised as true, false, or partially false due to a lack of evidence about its claims. This category includes articles in dispute and unproven articles.</p> </li> </ul> <p><strong>Cross-Lingual Task (German)</strong></p> <p>Along with the multi-class task for the English language, we have introduced a task for low-resourced language. We will provide the data for the test in the German language. The idea of the task is to use the English data and the concept of transfer to build a classification model for the German language.</p> <p><strong>Input Data</strong></p> <p>The data will be provided in the format of Id, title, text, rating, the domain; the description of the columns is as follows:</p> <ul> <li>ID- Unique identifier of the news article</li> <li>Title- Title of the news article</li> <li>text- Text mentioned inside the news article</li> <li>our rating - class of the news article as false, partially false, true, other</li> </ul> <p><strong>Output data format</strong></p> <ul> <li>public_id- Unique identifier of the news article</li> <li>predicted_rating- predicted class</li> </ul> <p>Sample File</p> <pre><code>public_id, predicted_rating 1, false 2, true</code></pre> <p><strong>IMPORTANT! </strong></p> <ol> <li>We have used the data from 2010 to 2022, and the content of fake news is mixed up with several topics like elections, COVID-19 etc.</li> </ol> <p><strong>Baseline:</strong> For this task, we have created a baseline system.&nbsp;The baseline system can be found at&nbsp;<a href="https://zenodo.org/record/6362498">https://zenodo.org/record/6362498</a></p> <p><strong>Related Work</strong></p> <ul> <li>Shahi GK. AMUSED: An Annotation Framework of Multi-modal Social Media Data. arXiv preprint arXiv:2010.00502. 2020 Oct 1.<a href="https://arxiv.org/pdf/2010.00502.pdf">https://arxiv.org/pdf/2010.00502.pdf</a></li> <li>G. K. Shahi and D. Nandini, &ldquo;FakeCovid &ndash; a multilingual cross-domain fact check news dataset for covid-19,&rdquo; in workshop Proceedings of the 14th International AAAI Conference on Web and Social Media, 2020.&nbsp;<a href="http://workshop-proceedings.icwsm.org/abstract?id=2020_14">http://workshop-proceedings.icwsm.org/abstract?id=2020_14</a></li> <li>Shahi, G. K., Dirkson, A., &amp; Majchrzak, T. A. (2021). An exploratory study of covid-19 misinformation on twitter.&nbsp;<em>Online Social Networks and Media</em>,&nbsp;<em>22</em>, 100104. doi:&nbsp;<a href="https://dx.doi.org/10.1016%2Fj.osnem.2020.100104">10.1016/j.osnem.2020.100104</a></li> <li>Shahi, G. K., Stru&szlig;, J. M., &amp; Mandl, T. (2021). Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection.&nbsp;<em>Working Notes of CLEF</em>.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barr&oacute;n-Cedeno, A., M&iacute;guez, R., Shaar, S., ... &amp; Mandl, T. (2021, March). The CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news. In&nbsp;<em>European Conference on Information Retrieval</em>&nbsp;(pp. 639-649). Springer, Cham.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barr&oacute;n-Cede&ntilde;o, A., M&iacute;guez, R., Shaar, S., ... &amp; Kartal, Y. S. (2021, September). Overview of the CLEF&ndash;2021 CheckThat! Lab on Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News. In&nbsp;<em>International Conference of the Cross-Language Evaluation Forum for European Languages</em>&nbsp;(pp. 264-291). Springer, Cham.</li> </ul>

openMay 2022View details →
dryad36/100

Data from: Fake news? The impact of information mismatch in mating behaviour

<p>Multiple cues are often used for mate choice in complex environments, potentially entailing mismatches between different sources of information. We address the consequences thereof for receivers using the spider mite <em>Tetranychus urticae</em>, in which virgin females are highly valuable mates compared to mated females, given first male sperm precedence. Accordingly, males prefer virgins and distinguish them using cues from the females and/or that are present on the substrate. Whereas the former are more reliable, the latter may allow for a faster or more long-distance response. However, there can be mismatched information between cues as females move and/or mate. Here, we tested the consequences thereof by exposing males to mated or virgin females on patches previously impregnated with cues deposited by females of either mating status. Male mating attempts were solely affected by substrate cues while female acceptance and the number of mating events were independently affected by both cues. Copulation duration, in contrast, depended mainly on the mating status of the female, with the number of copulations and the total time spent mating being intermediate in environments with mismatched information. Ultimately, male survival costs mirrored male investment in mating. These results suggest that, in environments with mismatched information, the substrate cues left by females are instrumental for males to find their mates, but they can also lead to males paying survival costs without the associated benefit of mating effectively, or suffering reduced costs at the expense of losing effective mating opportunities. The benefit of using multiple cues will then hinge upon the frequency of information mismatch, which itself should vary with the dynamics of populations.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Multilingual Fake News Detection Dataset: Gujarati, Hindi, Marathi, and Telugu

<p>This dataset is designed to support research in fake news detection across four major Indian languages: Gujarati, Hindi, Marathi, and Telugu. The dataset includes a diverse set of news articles collected from various sources, each labeled as either 'fake' or 'real'. The primary goal is to provide a resource that helps in the development and evaluation of natural language processing (NLP) models capable of detecting fake news in these regional languages.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

German Fake News Dataset "GermanFakeNC"

<p>&quot;GermanFakeNC&quot; is a German Fake News Corpus including 490 texts which were retrieved from German alternative online media sources. Every fake statement in the text was verified claim-by-claim by authoritative sources (e.g. from local police authorities, scientific studies, the police press office, etc.). The time interval for most of the news is established from December 2015 to March 2018.</p> <p>Steps to reproduce the data are described in the README file.</p> <p>Please cite:</p> <p>&nbsp;</p> <pre>@inproceedings{TPDL_Vogel19, author = {Inna Vogel and Peter Jiang}, title = {Fake News Detection with the New German Dataset &quot;GermanFakeNC&quot;}, booktitle = {Digital Libraries for Open Knowledge - 23rd International Conference on Theory and Practice of Digital Libraries, {TPDL} 2019, Oslo, Norway, September 9-12, 2019, Proceedings}, pages = {288--295}, year = {2019}, url = {https://doi.org/10.1007/978-3-030-30760-8\_25}, doi = {10.1007/978-3-030-30760-8\_25}, }</pre>

opencc-by-4.0Aug 2019View details →
zenodo36/100

Sample - Donald Trump's tweets mentioning the term 'fake news' (2017-2021)

<p>This is a dataset of the sample used for the paper &quot;<strong>Disintermediation and disinformation as a political strategy: using AI to analyse Trump&rsquo;s &ldquo;fake news&rdquo; discourse on Twitter&quot; </strong>by Diez-Gracia, A., S&aacute;nchez-Garc&iacute;a, P. &amp; Mart&iacute;n-Rom&aacute;n, J. (2023) published in the Journal &#39;El Profesional de la Informaci&oacute;n&#39; (pending DOI). Obtained by filtering the open-source database&nbsp;http://www.thetrumparchive.com</p> <p>Article full reference:&nbsp;Diez-Gracia, A., S&aacute;nchez-Garc&iacute;a, P. &amp; Mart&iacute;n-Rom&aacute;n, J. (2023). Disintermediation and disinformation as a political strategy: using AI to analyse Trump&#39;s &quot;fake news&quot; discourse on Twitter. El profesional de la informaci&oacute;n, vol., n.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Fake News Dataset with In-content Annotations and Detailed Lying Excerpts

<p>This dataset contains 95 fake news collected from two Brazilian fact-checking services (E-farsas and Boatos). We carefully annotated some excerpts of the fake news to enable deeper analyses of their falsehood. From these annotations, we divided the news into three groups: real news, totally fake news and fake news but which contain only a few lying snippets.</p><p>The annotated fragments were grouped based on four categories of falsehood: untrue, unverifiable fact, incorrectly named entity, and exaggeration. Other parts of the fake news, such as the passages around the fragments, were separated to add more information to the results of this research.</p>

opencc-by-4.0Nov 2020View details →
dryad36/100

Data from: Fake news? The impact of information mismatch in mating behaviour

Open the record for dataset details and reuse information.

publicMay 2024View details →
zenodo32/100

KEANE (faKe nEws At seNtence lEvel) dataset.

<p>This dataset is aimed at developing systems for detecting fake news. It contains links to articles, posts and other types of publications that have been previously evaluated by verifying entities. Sentence-level check-worthiness and fact-checking annotations are provided for each of these news items.</p>

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record