Skip to main content
zenodorestricted

TRACES Bulgarian Twitter Dataset on Covid-19 Annotated with Linguistic Markers of Lies

<p><strong>This dataset has been created within Project TRACES (more information: https://traces.gate-ai.eu/). The dataset contains 61411 tweet IDs of tweets, written in Bulgarian, with annotations. The dataset can be used for general use or for building lies and disinformation detection applications.&nbsp;</strong></p> <p><strong>Note: this dataset is not fact-checked, the social media messages have been retrieved via keywords. For fact-checked datasets, see our other datasets.</strong></p> <p><strong>The tweets (written between 1 Jan 2020 and 28&nbsp;June&nbsp;2022) have been collected via Twitter API under academic access in June 2022&nbsp;with the following keywords:</strong></p> <ul> <li> <p><strong>(Covid OR коронавирус OR Covid19 OR Covid-19 OR Covid_19) - without replies and without retweets</strong></p> </li> <li> <p><strong>(Корона OR корона OR Corona OR пандемия OR пандемията OR Spikevax OR SARS-CoV-2 OR бустерна доза) - with replies, but without retweets</strong></p> </li> </ul> <p><strong>Explanations of which fields can be used as markers of lies (or of intentional disinformation) are provided in our forthcoming paper (please cite it when using this dataset):&nbsp;</strong></p> <p>Irina Temnikova, Silvia Gargova, Ruslana Margova, Veneta Kireva, Ivo Dzhumerov, Tsvetelina Stefanova and Hristiana Nikolaeva&nbsp;(2023)&nbsp;New Bulgarian Resources for Detecting Disinformation.&nbsp;10th Language and Technology<br> Conference: Human Language Technologies as a Challenge for Computer Science and Linguistics (LTC&#39;23). Poznań. Poland.</p>

ShareScore

20/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
0
Reuse readiness
0
Engagement
8

Topics