Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

16

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

16 results for “fact checking”

Learn how ShareScore rates datasets ↗
zenodo48/100

Dataset. Responses to digital disinformation as part of hybrid threats: an evidence-based analysis on the effects of disinformation and the effectiveness of fact-checking/debunking

<p>Dataset&nbsp;from the meta-analysis carried out in the article Responses to digital disinformation as part of hybrid threats: a systematic review on the effects of disinformation and the effectiveness of fact-checking / debunking using the EU-HYBNET Meta-Analysis Survey Instrument for Evaluating the Effects of Disinformation and the Effectiveness of counter-responses</p>

opencc-by-4.0Apr 2021View details →
zenodo40/100

FaCov Dataset: COVID-19 Viral News and Rumors Fact-Check Articles Dataset

<p>The data were collected by web-scraping pages from the websites collected earlier, using the <a href="https://webscraper.io/">Web Scraper browser extension</a>.</p> <p>More specifically, the sections of these websites that dealt exclusively with COVID-19 related content were scraped. In cases where the website did not have such a specified section, the search functionality within the website was used to query terms related to COVID-19 and the articles in the search results were scraped. Also in some cases, all articles were scraped and those unrelated to COVID-19 were filtered out in the pre-processing stage. All the samples collected were then put together into one CSV</p> <p>The following information was extracted along with the articles:</p> <p>Title of the fact check article</p> <p>URL of the fact check article</p> <p>Claim being discussed in the article (if available)</p> <p>Summary of the fact check article (if available)</p> <p>Content of the fact check article&bull; Label assigned by the article to the claim</p> <p>Author of the fact check article (if available)</p> <p>Date of publication of the article (if available)</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and Models

<p>We propose the end-to-end multimodal fact-checking and explanation generation, where the input is a claim and a large collection of web sources, including articles, images, videos, and tweets, and the goal is to assess the truthfulness of the claim by retrieving relevant evidences and predicting a truthfulness label (i.e., support, refute and not enough information), and to generate a rationalization statement to explain the reasoning and ruling process. To support this research, we construct MOCHEG, a large-scale dataset consisting of 21,184 &nbsp;claims where each claim is annotated with a truthfulness label and ruling statement, with 43,148 text evidences and 15,373 image evidences. &nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Assessing the consistency of fact-checking in political debates

<p>Dataset of the mixed-method study named &quot;Assessing the consistency of fact-checking in political debates&quot;</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

FakeCovid- A Multilingual Cross domain Fact Check Dataset for COVID-19

<p>FakeCovid is the first multilingual cross-domain dataset of 7623 fact-checked news articles for COVID-19, collected from 04/01/2020 to 01/07/2020. We have collected the fact-checked articles from 92 fact-checking websites after obtaining references from Poynter and Snopes. We have manually annotated the collected articles into 11 categories of the fact-checked news according to their content. We&nbsp;ultimately generated dataset is in 40 languages from 105 countries.&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

Check Mate: Prioritizing User Generated Multi-Media Content for Fact-Checking

<p>Given volume of content and misinformation on social media, there is a need for systems that can support fact checkers by prioritizing content that needs to be fact checked. Prior research on prioritizing content for fact-checking has focused on news media articles, predominantly in English language. But there is an increasing amount of misinformation in user-generated content. Furthermore, misinformation is generated through information across modalities. In this paper we present a novel dataset that can be used to prioritize check-worthy posts from multi-media content in Hindi. It is unique in its 1) focus on user generated content, 2) multi-modality and 3) Hindi as the primary language of content. In addition, we also provide metadata for each post such as number of shares and likes of the post on ShareChat, a popular Indian social media platform, that allows for correlative analysis around virality and misinformation.&nbsp;</p>

opencc-by-4.0Sep 2020View details →
zenodo36/100

Repository of fact-checking websites and resources to combat climate mis/disinformation

<p>This dataset is the result of collaborative work for Deliverable 1.3 (WP1; T1.3) of the AGORA project. It compiles a list of fact-checking websites and resources dedicated to debunking climate change mis/disinformation. The identification of these resources was achieved by leveraging the expertise of our consortium and, therefore, most of the resources listed are in English, German, Italian and Spanish.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Datasets for the paper: Lost in Translation: Using Global Fact-Checks to Measure Multilingual Misinformation Prevalence, Spread, and Evolution

<p>FullData.csv.gz: Contains links to all claims in the data-set.</p> <ul> <li>publishing_date: Date on which the fact-check was published.</li> <li>claim_date: Date that claim was made.</li> <li>verdict: Rating given by the fact-checking organisation.</li> <li>language: Language of the claim.</li> <li>cluster_{threshold}: ID of the cluster that claim belongs to at all given clusters. Entry "0" means that claim is singleton and not clustered with any other claims.</li> </ul> <p>Embeddings.npy: Contains a dictionary linking each claim to it's embedding calculated with LaBSE.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

ClaimsKG - A Knowledge Graph of Fact-Checked Claims

<blockquote> <p><strong>The latest release of ClaimsKG is available in <a href="https://data.gesis.org/sharing/#!Detail/10.7802/2469">Datorium</a>.</strong></p> </blockquote> <p><strong>ClaimsKG</strong> is a <em>knowledge graph</em> of metadata information for thousands of fact-checked claims which facilitates structured queries about their truth values, authors, dates, and other kinds of metadata. <strong>ClaimsKG </strong>is generated through a (semi-)automated pipeline, which harvests claim-related data from popular fact-checking web sites, annotates them with related entities from DBpedia, and lifts all data to RDF using an RDF/S model that makes use of established vocabularies (such as schema.org).&nbsp;</p> <p><strong>ClaimsKG </strong>does NOT contain the text of the reviews from the fact-checking web sites;&nbsp;it only contains structured metadata information&nbsp;and links to the reviews.</p> <p>More information, such as statistics, query examples and&nbsp;a user friendly interface to explore the knowledge graph,&nbsp;is available at:&nbsp;<a href="https://data.gesis.org/claimskg/site">https://data.gesis.org/claimskg/site</a></p> <p>If you use <strong>ClaimsKG</strong>, please cite the below paper:</p> <p>Tchechmedjiev, Andon, Pavlos Fafalios, Katarina Boland, Malo Gasquet, Matth&auml;us Zloch, Benjamin Zapilko, Stefan Dietze, and Konstantin Todorov. &quot;ClaimsKG: a Knowledge Graph of Fact-Checked Claims.&quot; In&nbsp;<em>International Semantic Web Conference</em>, pp. 309-324. Springer, Cham, 2019.&nbsp;<a href="https://doi.org/10.1007/978-3-030-30796-7_20">https://doi.org/10.1007/978-3-030-30796-7_20</a><br> [<a href="http://users.ics.forth.gr/~fafalios/files/pubs/ISWC2019_ClaimsKG.pdf">pdf</a>, <a href="http://users.ics.forth.gr/~fafalios/files/bibs/ISWC2019_ClaimsKG.bib">bib</a>]&nbsp;</p>

opencc-by-nc-sa-4.0Sep 2019View details →
zenodo32/100

A Dataset of Fact-Checked Images Shared on WhatsApp during the Brazilian and Indian Elections

<p>(Dataset paper).&nbsp;In&nbsp;<em>Proceedings of the Int&#39;l AAAI Conference on Weblogs and Social Media (ICWSM&rsquo;20).&nbsp;</em>Atlanta, Georgia, U.S.&nbsp;June 2020.</p> <p>Abstract:&nbsp;Recently, messaging applications, such as WhatsApp, have been reportedly abused by misinformation campaigns, especially in Brazil and India. A notable form of abuse in WhatsApp relies on several manipulated images and memes containing all kinds of fake stories.&nbsp;In this work, we performed an extensive data collection from a large set of WhatsApp publicly accessible groups and &nbsp;fact-checking agency websites. This paper opens a novel dataset to the research community containing fact-checked fake images shared through WhatsApp for two distinct scenarios known for the &nbsp;spread of fake news on the platform: the 2018 Brazilian elections and the 2019 Indian elections.&nbsp;</p>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Fact checks versus problematic content in search rankings: SEO effects and the question of Google's content moderation

<p>Dataset released as supplementary material to the following publication:</p> <p>Kamila Koronska and Richard Rogers.2024. <span>Fact-checks versus problematic content in search rankings: SEO </span><span>effects and the question of Google&rsquo;s content moderation.Kamila Koronska and Richard Rogers. 2024. In ACM Web Science Conference (Websci &rsquo;24), May 21&ndash;24, 2024,&nbsp;Stuttgart, Germany. ACM, New York, NY, USA, 11 pages, https://<span>doi</span>.org/10.1145/3614419.3644017</span></p>

opencc-by-4.0Feb 2024View details →
zenodo32/100

Fact-checking Statistical Claims with Tables

<p>The surge of misinformation poses a serious problem for fact-checkers. Several initiatives for manualfact-checking have stepped up to combat this ordeal. However, computational methods are needed tomake the verification faster and keep up with the increasing abundance of false information. MachineLearning (ML) approaches have been proposed as a tool to ease the work of manual fact-checking.Specifically, the act of checking textual claims by using relational datasets has recently gained a lot oftraction. However, despite the abundance of proposed solutions, there has not been any formal definitionof the problem, nor a comparison across the different assumptions and results. In this work, we makea first attempt at solving these ambiguities. First, we formalize the problem by providing a generaldefinition that is applicable to all systems and that is agnostic to their assumptions. Second, we definegeneral dimensions to characterize different prominent systems in terms of assumptions and features.Finally, we report experimental results over three scenarios with corpora of real-world textual claims.</p>

opencc-by-4.0Jul 2021View details →
ClinicalTrials.gov32/100

Impact of a Mock-up Fact-checking Extension on HPV Vaccine Misinformation: A Survey Experiment

ClinicalTrials.gov study NCT06405048. IPD Sharing: YES. Countries: 1. Publications: 17.

controlledIPD-YESFeb 2026View details →
zenodo28/100

FactDrill: A Data Repository of Fact-checked Social Media Content to Study Fake News Incidents in India

<p>A dataset containing 22,435 fact-checked social media content to study fake news incidents in India. The dataset comprises news stories from 2013 to the year 2020, covering 13 different languages spoken in the country. There are&nbsp;14 different attributes present in the dataset.</p>

openJan 2022View details →
zenodo28/100

Temporal Fact Checking dataset

<p>This is a temporal fact checking dataset. Published for reviewers</p>

opencc-by-4.0Apr 2023View details →
zenodo12/100

Dataset for the paper: "Monant Medical Misinformation Dataset: Mapping Articles to Fact-Checked Claims"

<p><strong>Overview</strong></p> <p>This dataset of medical misinformation was collected and is published by <a href="https://kinit.sk/">Kempelen Institute of Intelligent Technologies (KInIT)</a>.&nbsp;It consists of approx. 317k news articles and blog posts on medical topics published between January 1, 1998 and February 1, 2022 from a total of 207 reliable and unreliable sources. The dataset contains full-texts of the articles, their original source URL and other extracted metadata. If a source has a credibility score available (e.g., from Media Bias/Fact Check), it is also included in the form of annotation.&nbsp;Besides the articles, the dataset contains around 3.5k fact-checks and extracted verified medical claims with their unified veracity ratings published by fact-checking organisations such as Snopes or FullFact. Lastly and most importantly, the dataset contains 573 manually and more than 51k automatically labelled mappings between previously verified claims and the articles; mappings consist of two values: <em>claim presence </em>(i.e., whether a claim is contained in the given article) and <em>article stance </em>(i.e., whether the given article supports or rejects the claim or provides both sides of the argument).</p> <p>The dataset is primarily intended to be used as a training and evaluation set for machine learning methods for claim presence detection and article stance classification, but it enables a range of other misinformation related tasks, such as misinformation characterisation or analyses of misinformation spreading.</p> <p>Its novelty and our main contributions lie in (1) focus on medical news article and blog posts as opposed to social media posts or political discussions; (2) providing multiple modalities (beside full-texts of the articles, there are also images and videos), thus enabling research of multimodal approaches; (3) mapping of the articles to the fact-checked claims (with manual as well as predicted labels); (4) providing source credibility labels for 95% of all articles and other potential sources of weak labels that can be mined from the articles' content and metadata.</p> <p>The dataset is associated with the research paper "Monant Medical Misinformation Dataset: Mapping Articles to Fact-Checked Claims" accepted and presented at ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '22).&nbsp;</p> <p>The accompanying&nbsp;<a href="https://github.com/kinit-sk/medical-misinformation-dataset">Github repository</a> provides a small static sample of the dataset and the dataset's descriptive analysis in a form of Jupyter notebooks.</p> <p>In order to obtain an access to the full dataset (in the CSV format), please, request the access by following the instructions provided below.</p> <p>&nbsp;</p> <p><strong>Note: </strong>Please, check also our <a href="https://doi.org/10.5281/zenodo.7737982">MultiClaim Dataset</a> that provides a more recent, a larger, and a highly multilingual dataset of fact-checked claims, social media posts and relations between them.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>If you use this dataset in any publication, project, tool or in any other form, please, cite the following papers:</p> <pre><code>@inproceedings{SrbaMonantPlatform, author = {Srba, Ivan and Moro, Robert and Simko, Jakub and Sevcech, Jakub and Chuda, Daniela and Navrat, Pavol and Bielikova, Maria}, booktitle = {Proceedings of Workshop on Reducing Online Misinformation Exposure (ROME 2019)}, pages = {1--7}, title = {Monant: Universal and Extensible Platform for Monitoring, Detection and Mitigation of Antisocial Behavior}, year = {2019} }</code></pre> <pre><code>@inproceedings{SrbaMonantMedicalDataset, author = {Srba, Ivan and Pecher, Branislav and Tomlein Matus and Moro, Robert and Stefancova, Elena and Simko, Jakub and Bielikova, Maria}, booktitle = {Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '22)}, numpages = {11}, title = {Monant Medical Misinformation Dataset: Mapping Articles to Fact-Checked Claims}, year = {2022}, doi = {10.1145/3477495.3531726}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3477495.3531726}, } </code></pre> <p><br><strong>Dataset creation process</strong></p> <p>In order to create this dataset (and to continuously obtain new data), we used our research platform <a href="https://rome2019.github.io/papers/Srba_etal_ROME2019.pdf">Monant</a>. The Monant platform provides so called data providers to extract news articles/blogs from news/blog sites as well as fact-checking articles from fact-checking sites. General parsers (from RSS feeds, Wordpress sites, Google Fact Check Tool, etc.) as well as custom crawler and parsers were implemented (e.g., for fact checking site Snopes.com). All data is stored in the unified format in a central data storage.<br><strong>&nbsp;&nbsp; &nbsp;<br>&nbsp; &nbsp;<br>Ethical considerations</strong></p> <p>The dataset was collected and is published for research purposes only. We collected only publicly available content of news/blog articles. The dataset contains identities of authors of the articles if they were stated in the original source; we left this information, since the presence of an author's name can be a strong credibility indicator. However, we anonymised the identities of the authors of discussion posts included in the dataset.&nbsp;</p> <p>The main identified ethical issue related to the presented dataset lies in the risk of mislabelling of an article as supporting a false fact-checked claim and, to a lesser extent, in mislabelling an article as not containing a false claim or not supporting it when it actually does. To minimise these risks, we developed a labelling methodology and require an agreement of at least two independent annotators to assign a claim presence or article stance label to an article. It is also worth noting that we do not label an article as a whole as false or true. Nevertheless, we provide partial article-claim pair veracities based on the combination of claim presence and article stance labels.</p> <p>As to the veracity labels of the fact-checked claims and the credibility (reliability) labels of the articles' sources, we take these from the fact-checking sites and external listings such as Media Bias/Fact Check as they are and refer to their methodologies for more details on how they were established.</p> <p>Lastly, the dataset also contains automatically predicted labels of claim presence and article stance using our baselines described in the next section. These methods have their limitations and work with certain accuracy as reported in this paper. This should be taken into account when interpreting them.&nbsp;<br><strong>&nbsp; &nbsp;<br>&nbsp; &nbsp;<br>Reporting mistakes in the dataset</strong><br>The mean to report considerable mistakes in raw collected data or in manual annotations is by creating a new issue in the accompanying&nbsp;<a href="https://github.com/kinit-sk/medical-misinformation-dataset">Github repository</a>.&nbsp;Alternately, general enquiries or&nbsp;requests can be sent at info [at] kinit.sk.</p> <p><br><strong>Dataset structure</strong></p> <p><em><strong>Raw data</strong></em></p> <p>At first, the dataset contains so called raw data (i.e., data extracted by the Web monitoring module of Monant platform and stored in exactly the same form as they appear at the original websites). Raw data consist of articles from news sites and blogs (e.g. naturalnews.com), discussions attached to such articles, fact-checking articles from fact-checking portals (e.g. snopes.com). In addition, the dataset contains feedback (number of likes, shares, comments) provided by user on social network Facebook which is regularly extracted for all news/blogs articles.</p> <p>Raw data are contained in these CSV files:</p> <ul> <li>sources.csv</li> <li>articles.csv</li> <li>article_media.csv</li> <li>article_authors.csv</li> <li>discussion_posts.csv</li> <li>discussion_post_authors.csv</li> <li>fact_checking_articles.csv</li> <li>fact_checking_article_media.csv</li> <li>claims.csv</li> <li>feedback_facebook.csv</li> </ul> <p><em>Note: Personal information about discussion posts' authors (name, website, gravatar) are anonymised.</em></p> <p><br><em><strong>Annotations</strong></em></p> <p>Secondly, the dataset contains so called annotations. Entity annotations describe the individual raw data entities (e.g., article, source). Relation annotations describe relation between two of such entities.</p> <p>Each annotation is described by the following attributes:</p> <ol> <li>category of annotation (`annotation_category`). Possible values: label (annotation corresponds to ground truth, determined by human experts) and prediction (annotation was created by means of AI method).</li> <li>type of annotation (`annotation_type_id`). Example values: Source reliability (binary), Claim presence. The list of possible values can be obtained from enumeration in annotation_types.csv.</li> <li>method which created annotation (`method_id`). Example values: Expert-based source reliability evaluation, Fact-checking article to claim transformation method. The list of possible values can be obtained from enumeration methods.csv.</li> <li>its value (`value`). The value is stored in JSON format and its structure differs according to particular annotation type.</li> </ol> <p><br>At the same time, annotations are associated with a particular object identified by:</p> <ol> <li>entity type (parameter `entity_type` in case of entity annotations, or `source_entity_type` and `target_entity_type` in case of relation annotations). Possible values: sources, articles, fact-checking-articles.</li> <li>entity id (parameter `entity_id` in case of entity annotations, or `source_entity_id` and `target_entity_id` in case of relation annotations).</li> </ol> <p><br>The dataset provides specifically these entity annotations:</p> <ul> <li>Source reliability (binary). Determines validity of source (website) at a binary scale with two options: reliable source and unreliable source.</li> <li>Article veracity. Aggregated information about veracity from article-claim pairs.</li> </ul> <p>The dataset provides specifically these relation annotations:</p> <ul> <li>Fact-checking article to claim mapping. Determines mapping between fact-checking article and claim.</li> <li>Claim presence. Determines presence of claim in article.</li> <li>Claim stance. Determines stance of an article to a claim.&nbsp;</li> </ul> <p><br>Annotations are contained in these CSV files:</p> <ul> <li>entity_annotations.csv</li> <li>relation_annotations.csv</li> </ul> <p><em>Note: Identification of human annotators authors (email provided in the annotation app) is anonymised.</em></p> <p>&nbsp;</p> <p><em><strong>Enumerations</strong></em></p> <p>Finally, the dataset provides additional CSV files with enumerations:</p> <ul> <li>media_types.csv</li> <li>source_types.csv</li> <li>annotation_types.csv</li> <li>methods.csv</li> </ul>

restrictedFeb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record