Turkish Political Fake News Dataset (TPFND)
<div> <div> <div> <div> <div> <div> <div> <div> <p>The Turkish Political Fake News Dataset (TPFND) presented in this repository was developed as part of the doctoral dissertation.</p> <p>This repository contains the following components:</p> <ol> <li> <p><strong>Original TPFND</strong>: The Turkish Political Fake News Dataset (TPFND) is a meticulously curated dataset developed to address the scarcity of high-quality Turkish political fake news data. The dataset reflects the real-world class imbalance, with 2,308 instances (25%) of verified political fake news and 6,922 instances (75%) of factual political news. The classification of fake news is grounded in evidence from reputable Turkish fact-checking organizations, ensuring the dataset's credibility and reliability.</p> </li> <li> <p><strong>Augmented TPFND</strong>: To mitigate the inherent class imbalance in the original TPFND, this repository includes an augmented version of the dataset. The augmentation process employed a large language model (LLM), specifically the Turkish LLaMA-3 8B model, to generate synthetic samples of Turkish political fake news. This approach effectively increased the representation of the minority class, raising the proportion of fake news instances from 25% to 40% within the augmented dataset.</p> </li> <li> <p><strong>Synthetic Data</strong>: The synthetic data generated by the LLM is also provided in this repository. This synthetic data was carefully calibrated to closely mimic the characteristics and nuances of the original fake news samples, ensuring the integrity and authenticity of the augmented dataset.</p> </li> </ol> <p>The TPFND and its augmented version, along with the synthetic data, serve as valuable resources for researchers and practitioners working on fake news detection, particularly in the Turkish context. The dataset's rigorous development, the incorporation of LLM-generated synthetic data, and the addressing of class imbalance issues make this repository a significant contribution to the field of Turkish political fake news research.</p> <p>Researchers can leverage this dataset and the accompanying synthetic data to develop and evaluate advanced machine learning models for fake news detection. The availability of a controlled, high-quality dataset, coupled with the synthetic data, enables the exploration of novel techniques to enhance the accuracy and robustness of fake news detection, ultimately supporting the fight against the spread of fake news in the digital age.</p> </div> </div> </div> </div> </div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <div> <div> </div> </div> </div> </div> </div> </div> </div> </div>
ShareScore
20/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 8
- Reuse readiness
- 0
- Engagement
- 0