Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

8

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

8 results for “Fake detection”

Learn how ShareScore rates datasets ↗
zenodo40/100

Datasets from Costa Rican news sources for fake news detection

<p>Today, technology has changed the way information is propagated and how the message is received. The interpretation of the news may have different angles depending on the source of origin. Because of this, there has been an increase in misinformation, in the way of influencing public opinion and in how we perceive or estimate reality.<br> The objective of this beta dataset is to be used for the evaluation of data mining models that allow the classification of true or potentially fake news that are generated by Costa Rican news sites only.&nbsp; This is intended to assess the level of reliability of the models and extend the scope of this research in future work.</p> <p>The dataset has been pre-processed (standarized using lower cases, lemmatized and removed any possible noise from it) and analyzed using LIWC dictionaries. One version has the news text in Spanish and was processed using LIWC2007 dictionary in Spanish. The second version was processed using LIWC2015 English dictionary and has the news text in English. The reason to having two versions is to be able to test using the newer LIWC dictionary which includes more Summary Language Variables that the Spanish version doesn&#39;t have and analyze how this and other variables can contribute to different results when creating models.&nbsp;</p> <p>The file &quot;DescripcionVariables&quot; provides a description of all variables used.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

CFAD: A Chinese Dataset for Fake Audio Detection

<p>Fake audio detection is a growing concern and some relevant datasets have been designed for research. However, there is no standard public Chinese dataset under complex conditions.</p> <p>In this paper, we aim to fill in the gap and design a Chinese fake audio detection dataset (CFAD) for studying more generalized detection methods. Twelve mainstream speech-generation techniques are used to generate fake audio. To simulate the real-life scenarios, three noise datasets are selected for noise adding at five different signal-to-noise ratios, and six codecs are considered for audio transcoding. CFAD dataset can be used not only for fake audio detection but also for detecting the algorithms of fake utterances for audio forensics. Baseline results are presented with analysis. The results that show fake audio detection methods with generalization remain challenging. The CFAD dataset is publicly available&nbsp;on GitHub&nbsp;https://github.com/ADDchallenge/CFAD</p> <p>&nbsp;</p> <p>CFAD dataset considers 12 types of fake audio, 11 of which are generated by different speech synthesis techniques and the remaining one is partially fake type. Partially fake audio is completely different from synthesis speech and thus can better evaluate the generalization of the detection model to unknown types. The real audio is collected from 6 different corpora to increase the diversity of real category distributions, which makes model less prone to artifact from a single database. For robustness evaluation, we additionally simulate background noise and media codecs that might occur in real life and provide detailed labels, including fake type, real source, noise type, signal noise ratio (SNR), and media codecs. Overall, CFAD dataset consists of three different versions, named clean, noisy, and codec versions.</p> <p>Each version of the dataset is divided into disjoint training, development, and test sets in the same way. There is no speaker overlap across these three subsets. Each test set is further divided into seen and unseen test sets. Unseen test sets can evaluate the generalization of the methods to unknown types. It is worth mentioning that both real audio and fake audio in the unseen test set are unknown to the model.</p> <p>For the noisy speech part, we select three noise databases for simulation. Additive noises are added to each audio in the clean dataset at 5 different SNRs. The additive noises of the unseen test set and the remaining subsets come from different noise databases.</p> <p>For the codec speech part, we select six different codecs. Two of them are applied for unseen test set.</p> <p>In each version (clean, noisy, and codec versions) of the CFAD dataset, there are 138400 utterances in training set, 14400 utterances in development set, 42000 utterances in seen test set, and 21000 utterances in unseen test set.</p> <p>&nbsp;</p> <p>Clean Real Audios Collection</p> <p>From the point of eliminating the interference of irrelevant factors, we collect clean real audios from&nbsp;two aspects: 5 open resources from OpenSLR platform (http://www.openslr.org/12/) and one self-recording dataset.&nbsp;</p> <p>&nbsp;</p> <p>Clean Fake Audios Generation</p> <p>We select 11 representative speech synthesis methods to generate the fake audios and one partially fake audios.</p> <p>&nbsp;</p> <p>Noisy Audios Simulation</p> <p>Noisy audios aim to quantify the robustness of the methods under noisy conditions. To simulate the real-life scenarios, we artificially sample the noise signals and add them to clean audios at 5 different&nbsp;SNRs, which are 0dB, 5dB, 10dB, 15dB and 20dB. Additive noises are selected from three noise databases: PNL 100 Nonspeech Sounds, NOISEX-92, and TAU Urban Acoustic Scenes.</p> <p>&nbsp;</p> <p>Audio Transcoding</p> <p>The Codec version aims to quantify the robustness of the methods under different format conversions. We select a total of six codecs. For the training, development, and seen test sets in codec version, mp3, flac, ogg, and m4a are used. For the unseen test set of the codec version, aac, and wma are used.</p> <p>Audio transcoding operation is operated on the audio in the clean version. Each clean audio will be randomly transformed with one of the candidate codecs and converted back to original WAV files using ffmpeg toolkits.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>This data set is licensed with a CC BY-NC-ND 4.0 license.</p> <p>You can cite the data using the following BibTeX entry.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

COVID-19 Fake News Detection Dataset

<p>Please note that this data set was originally shared by Patwa et al. (2021) on GitHub.&nbsp;</p> <p><strong>Reference </strong></p> <p>Patwa, P., Sharma, S. Pykl, S., Guptha, V., Kumari, G., Akhtar, M. S., Ekbal, A., Das A. &amp; Chakraborty, T. (2021). Fighting an Infodemic: COVID-19 Fake News Dataset. Combating Online Hostile Posts in Regional Languages during Emergency Situation, Cham, Springer International Publishing. https://doi.org/10.1007/978-3-030-73696-5_3</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

CT-FAN: A Multilingual dataset for Fake News Detection

<p><strong>By downloading the data, you agree with the terms &amp; conditions mentioned below:</strong></p> <p><strong>Data Access:&nbsp;</strong>The data in the research collection may only be used for research purposes. Portions of the data are copyrighted and have commercial value as data, so you must be careful to use them only for research purposes.&nbsp;</p> <p>Summaries, analyses and interpretations of the linguistic properties of the information may be derived and published, provided it is impossible to reconstruct the information from these summaries.&nbsp;You may not try identifying the individuals whose texts are included in this dataset. You may not try to identify the original entry on the fact-checking site. You are not permitted to publish any portion of the dataset besides summary statistics or share it with anyone else.</p> <p>We grant you the right to access the collection&#39;s content as described in this agreement. You may not otherwise make unauthorised commercial use of, reproduce, prepare derivative works, distribute copies, perform, or publicly display the collection or parts of it. You are responsible for keeping and storing the data in a way that others cannot access. The data is provided free of charge.</p> <p><strong>Citation</strong></p> <p>Please cite our work&nbsp;as</p> <pre>@InProceedings{clef-checkthat:2022:task3, author = {K{\&quot;o}hler, Juliane and Shahi, Gautam Kishore and Stru{\ss}, Julia Maria and Wiegand, Michael and Siegel, Melanie and Mandl, Thomas}, title = &quot;Overview of the {CLEF}-2022 {CheckThat}! Lab Task 3 on Fake News Detection&quot;, year = {2022}, booktitle = &quot;Working Notes of CLEF 2022---Conference and Labs of the Evaluation Forum&quot;, series = {CLEF~&#39;2022}, address = {Bologna, Italy},} </pre> <pre>@article{shahi2021overview, title={Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection}, author={Shahi, Gautam Kishore and Stru{\ss}, Julia Maria and Mandl, Thomas}, journal={Working Notes of CLEF}, year={2021} }</pre> <p><strong>Problem Definition:</strong> Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other (e.g., claims in dispute) and detect the topical domain of the article. This task will run in <strong>English and German.</strong></p> <p><strong>Task 3:</strong> <strong>Multi-class fake news detection of news articles (English)</strong>&nbsp;Sub-task A would detect fake news designed as a four-class classification problem. Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other. The training data will be released in batches and roughly about 1264 articles with the respective label in English language. Our definitions for the categories are as follows:</p> <ul> <li> <p>False - The main claim made in an article is untrue.</p> </li> <li> <p>Partially False - The main claim of an article is a mixture of true and false information. The article contains partially true and partially false information but cannot be considered 100% true. It includes all articles in categories like partially false, partially true, mostly true, miscaptioned, misleading etc., as defined by different fact-checking services.</p> </li> <li> <p>True - This rating indicates that the primary elements of the main claim are demonstrably true.</p> </li> <li> <p>Other- An article that cannot be categorised as true, false, or partially false due to a lack of evidence about its claims. This category includes articles in dispute and unproven articles.</p> </li> </ul> <p><strong>Cross-Lingual Task (German)</strong></p> <p>Along with the multi-class task for the English language, we have introduced a task for low-resourced language. We will provide the data for the test in the German language. The idea of the task is to use the English data and the concept of transfer to build a classification model for the German language.</p> <p><strong>Input Data</strong></p> <p>The data will be provided in the format of Id, title, text, rating, the domain; the description of the columns is as follows:</p> <ul> <li>ID- Unique identifier of the news article</li> <li>Title- Title of the news article</li> <li>text- Text mentioned inside the news article</li> <li>our rating - class of the news article as false, partially false, true, other</li> </ul> <p><strong>Output data format</strong></p> <ul> <li>public_id- Unique identifier of the news article</li> <li>predicted_rating- predicted class</li> </ul> <p>Sample File</p> <pre><code>public_id, predicted_rating 1, false 2, true</code></pre> <p><strong>IMPORTANT! </strong></p> <ol> <li>We have used the data from 2010 to 2022, and the content of fake news is mixed up with several topics like elections, COVID-19 etc.</li> </ol> <p><strong>Baseline:</strong> For this task, we have created a baseline system.&nbsp;The baseline system can be found at&nbsp;<a href="https://zenodo.org/record/6362498">https://zenodo.org/record/6362498</a></p> <p><strong>Related Work</strong></p> <ul> <li>Shahi GK. AMUSED: An Annotation Framework of Multi-modal Social Media Data. arXiv preprint arXiv:2010.00502. 2020 Oct 1.<a href="https://arxiv.org/pdf/2010.00502.pdf">https://arxiv.org/pdf/2010.00502.pdf</a></li> <li>G. K. Shahi and D. Nandini, &ldquo;FakeCovid &ndash; a multilingual cross-domain fact check news dataset for covid-19,&rdquo; in workshop Proceedings of the 14th International AAAI Conference on Web and Social Media, 2020.&nbsp;<a href="http://workshop-proceedings.icwsm.org/abstract?id=2020_14">http://workshop-proceedings.icwsm.org/abstract?id=2020_14</a></li> <li>Shahi, G. K., Dirkson, A., &amp; Majchrzak, T. A. (2021). An exploratory study of covid-19 misinformation on twitter.&nbsp;<em>Online Social Networks and Media</em>,&nbsp;<em>22</em>, 100104. doi:&nbsp;<a href="https://dx.doi.org/10.1016%2Fj.osnem.2020.100104">10.1016/j.osnem.2020.100104</a></li> <li>Shahi, G. K., Stru&szlig;, J. M., &amp; Mandl, T. (2021). Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection.&nbsp;<em>Working Notes of CLEF</em>.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barr&oacute;n-Cedeno, A., M&iacute;guez, R., Shaar, S., ... &amp; Mandl, T. (2021, March). The CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news. In&nbsp;<em>European Conference on Information Retrieval</em>&nbsp;(pp. 639-649). Springer, Cham.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barr&oacute;n-Cede&ntilde;o, A., M&iacute;guez, R., Shaar, S., ... &amp; Kartal, Y. S. (2021, September). Overview of the CLEF&ndash;2021 CheckThat! Lab on Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News. In&nbsp;<em>International Conference of the Cross-Language Evaluation Forum for European Languages</em>&nbsp;(pp. 264-291). Springer, Cham.</li> </ul>

openMay 2022View details →
zenodo36/100

Multilingual Fake News Detection Dataset: Gujarati, Hindi, Marathi, and Telugu

<p>This dataset is designed to support research in fake news detection across four major Indian languages: Gujarati, Hindi, Marathi, and Telugu. The dataset includes a diverse set of news articles collected from various sources, each labeled as either 'fake' or 'real'. The primary goal is to provide a resource that helps in the development and evaluation of natural language processing (NLP) models capable of detecting fake news in these regional languages.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

An exploratory study (with and without time pressure) using mouse dynamics to detect faking-good behavior in the MMPI-2 and PPI-R validity scales

<p>Please find here&nbsp;the dataset&nbsp;generated and analyzed during the study entitled &quot;Can mouse dynamics detect faking-good behavior in personality questionnaires? An exploratory study (with and without time pressure) using the MMPI-2 and PPI-R validity scales&quot;.&nbsp;&nbsp;Moreover, here you can find the source code of the experiment to execute the task using MouseTracker software, the code of the statistical analysis and a file containing the instructions to replicate ML model results reported in the original paper.</p>

opencc-by-4.0Oct 2019View details →
zenodo28/100

New Outlooks on Backdoor Attacks for Fake Speech Detection

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo16/100

CT-FAN-22 corpus: A Multilingual dataset for Fake News Detection

<p><strong>Data Access:&nbsp;</strong>The data in the research collection provided&nbsp;may only be used for research purposes. Portions of the data are copyrighted and have commercial value as data, so you must be careful to use it only for research purposes. Due to these restrictions, the collection is not open data. Please download the Agreement at <a href="https://drive.google.com/file/d/1QU-rw4D26r3F04FB63hTvToxOvDKaKdv/view?usp=sharing">Data Sharing Agreement</a>&nbsp;and send the signed form to <a href="mailto:fakenewstask@gmail.com">fakenewstask@gmail.com</a> .</p> <p><strong>Citation</strong></p> <p>Please cite our work&nbsp;as</p> <pre>@article{shahi2021overview, title={Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection}, author={Shahi, Gautam Kishore and Stru{\ss}, Julia Maria and Mandl, Thomas}, journal={Working Notes of CLEF}, year={2021} }</pre> <p><strong>Problem Definition:</strong> Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other (e.g., claims in dispute) and detect the topical domain of the article. This task will run in <strong>English and German.</strong></p> <p><strong>Subtask 3:</strong> <strong>Multi-class fake news detection of news articles (English)</strong>&nbsp;Sub-task A would detect fake news designed as a four-class classification problem. The training data will be released in batches and roughly about 900 articles with the respective label. Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other. Our definitions for the categories are as follows:</p> <ul> <li> <p>False - The main claim made in an article is untrue.</p> </li> <li> <p>Partially False - The main claim of an article is a mixture of true and false information. The article contains partially true and partially false information but cannot be considered 100% true. It includes all articles in categories like partially false, partially true, mostly true, miscaptioned, misleading etc., as defined by different fact-checking services.</p> </li> <li> <p>True - This rating indicates that the primary elements of the main claim are demonstrably true.</p> </li> <li> <p>Other- An article that cannot be categorised as true, false, or partially false due to lack of evidence about its claims. This category includes articles in dispute and unproven articles.</p> </li> </ul> <p><strong>Input Data</strong></p> <p>The data will be provided in the format of Id, title, text, rating, the domain; the description of the columns is as follows:</p> <p><strong>Task 3</strong></p> <ul> <li>ID- Unique identifier of the news article</li> <li>Title- Title of the news article</li> <li>text- Text mentioned inside the news article</li> <li>our rating - class of the news article as false, partially false, true, other</li> </ul> <p><strong>Output data format</strong></p> <p><strong>Task 3</strong></p> <ul> <li>public_id- Unique identifier of the news article</li> <li>predicted_rating- predicted class</li> </ul> <p>Sample File</p> <pre><code>public_id, predicted_rating 1, false 2, true</code></pre> <p>Sample file</p> <pre><code>public_id, predicted_domain 1, health 2, crime</code></pre> <p><strong>Additional data for Training</strong></p> <p>To train your model, the participant can use additional data with a similar format; some datasets are available over the web. We don&#39;t provide the background truth for those datasets. For testing, we will not use any articles from other datasets. Some of the possible sources:</p> <ul> <li><a href="https://www.kaggle.com/liberoliber/onion-notonion-datasets">Fakenews Classification Datasets</a></li> <li><a href="https://www.kaggle.com/c/fakenewskdd2020/overview">Fake News Detection Challenge KDD 2020</a></li> <li><a href="https://www.kaggle.com/mdepak/fakenewsnet?select=PolitiFact_real_news_content.csv">FakeNewsNet</a></li> </ul> <p><strong>IMPORTANT! </strong></p> <ol> <li>We have used the data from 2010 to 2021, and the content of fake news is mixed up with several topics like election, COVID-19 etc.</li> </ol> <p><strong>Evaluation Metrics</strong></p> <p>This task is evaluated as a classification task. We will use the F1-macro measure for the ranking of teams. There is&nbsp;a limit of 5 runs&nbsp;(total and not per day), and only one person from a team is allowed to submit runs.</p> <p><strong>Submission Link:&nbsp;</strong><a href="https://codalab.org/">Coming soon</a></p> <p><strong>Related Work</strong></p> <ul> <li>Shahi, G. K., Stru&szlig;, J. M., &amp; Mandl, T. (2021). Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection.&nbsp;<em>Working Notes of CLEF</em>.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barr&oacute;n-Cede&ntilde;o, A., M&iacute;guez, R., Shaar, S., ... &amp; Mandl, T. (2021, March). The CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news. In&nbsp;<em>European Conference on Information Retrieval</em>&nbsp;(pp. 639-649). Springer, Cham.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barr&oacute;n-Cede&ntilde;o, A., M&iacute;guez, R., Shaar, S., ... &amp; Kartal, Y. S. (2021, September). Overview of the CLEF&ndash;2021 CheckThat! Lab on Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News. In&nbsp;<em>International Conference of the Cross-Language Evaluation Forum for European Languages</em>&nbsp;(pp. 264-291). Springer, Cham.</li> <li>Shahi GK. AMUSED: An Annotation Framework of Multi-modal Social Media Data. arXiv preprint arXiv:2010.00502. 2020 Oct 1.<a href="https://arxiv.org/pdf/2010.00502.pdf">https://arxiv.org/pdf/2010.00502.pdf</a></li> <li>G. K. Shahi and D. Nandini, &ldquo;FakeCovid &ndash; a multilingualcross-domain fact check news dataset for covid-19,&rdquo; inWorkshop Proceedings of the 14th International AAAIConference on Web and Social Media, 2020.&nbsp;<a href="http://workshop-proceedings.icwsm.org/abstract?id=2020_14">http://workshop-proceedings.icwsm.org/abstract?id=2020_14</a></li> <li>Shahi, G. K., Dirkson, A., &amp; Majchrzak, T. A. (2021). An exploratory study of covid-19 misinformation on twitter.&nbsp;<em>Online Social Networks and Media</em>,&nbsp;<em>22</em>, 100104. doi:&nbsp;<a href="https://dx.doi.org/10.1016%2Fj.osnem.2020.100104">10.1016/j.osnem.2020.100104</a></li> </ul>

restrictedApr 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record