TRACES Bulgarian Twitter Dataset on Lies and Manipulation Annotated with Linguistic Markers of Lies
<p><strong>This dataset has been created within Project TRACES (more information: https://traces.gate-ai.eu/). The dataset is in .csv format and contains 32518 tweet IDs of tweets, written in Bulgarian, with annotations. Note: this dataset is not fact-checked, the social media messages have been retrieved via keywords. For fact-checked datasets, see our other datasets.</strong></p> <p><strong>The dataset can be used for general purposes or for building lies and disinformation detection applications (by using the annotations with the linguistic markers of lies). </strong></p> <p><strong>The tweets (written between 1 Jan 2020-27 June 2022) have been collected via Twitter API under academic access in June-July 2022 with the following keywords:</strong></p> <p><strong>(лъжа OR лъжи OR лицемерие OR лъжат OR излъга OR измама OR измамници OR измами OR лъжец OR лъжци) </strong></p> <p><strong>(фалшиви OR fakenews OR невярно OR неверни OR подвеждащи OR подвеждащо OR неистини) - without retweets</strong></p> <p><strong>(манипулация OR манипулира OR стъкмистика OR крие OR далавераджия OR далавери OR далавера) - without retweets</strong></p> <p><strong>Explanation of which fields can be used as markers of lies (or of intentional disinformation) are provided in our forthcoming paper:</strong></p> <p>Irina Temnikova, Silvia Gargova, Ruslana Margova, Veneta Kireva, Ivo Dzhumerov, Tsvetelina Stefanova and Hristiana Nikolaeva (2023) New Bulgarian Resources for Detecting Disinformation. 10th Language and Technology<br> Conference: Human Language Technologies as a Challenge for Computer Science and Linguistics (LTC'23). Poznań. Poland.</p> <p> </p>
ShareScore
20/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 0
- Reuse readiness
- 0
- Engagement
- 8