Skip to main content
zenodorestricted

Comprehensive Collection of English, German, Russian and Ukrainian Tweets Containing the Word or Hashtag Ukraine During the Russian Invasion, February 2022 until May 2023

<p>Comprehensive dataset of Tweets containing the keyword 'ukraine' (in German, Russian and Ukrainian) as well as '#ukraine' (in English) since the Russian Ukraine Invasion in February 2022.&nbsp;The user handle column has been excluded to protect deleted accounts that have not been retweeted or replied to.&nbsp;Tweets have been collected via the Academic API using the Search endpoint in four languages:</p> <table> <tbody> <tr> <td><strong>Language</strong></td> <td><strong>Query</strong></td> <td><strong> Name in Dataset (column: event)&nbsp; </strong></td> <td><strong>Number of Tweets<br></strong></td> </tr> <tr> <td>English</td> <td>#ukraine AND lang='en'&nbsp;</td> <td>ukraine-en-hashtag</td> <td>45.8 million</td> </tr> <tr> <td>German</td> <td>ukraine AND lang='de'&nbsp;</td> <td>ukraine</td> <td>19.6 million</td> </tr> <tr> <td>Russian</td> <td>Украина AND lang:ru&nbsp;</td> <td>ukraine-ru</td> <td>5.02 million</td> </tr> <tr> <td>Ukrainian</td> <td>Україна AND lang:uk&nbsp;</td> <td>ukraine-uk</td> <td>4.1 million</td> </tr> </tbody> </table> <p><strong>Collection dates</strong></p> <p>Details on collection dates per Tweet (e.g. to compare with creation dates) as well as the IDs of Tweets for consistency checks can be found here: <a href="https://github.com/Leibniz-HBI/ukraine_twitter_data">https://github.com/Leibniz-HBI/ukraine_twitter_data</a>&nbsp;(<a href="https://doi.org/10.17605/OSF.IO/RTQXN">https://doi.org/10.17605/OSF.IO/RTQXN</a>)</p> <p><strong>File Naming Scheme</strong></p> <p>To enable downloads of selected timeframes and languages, the files are named by language, start and end date of the tweet creation timestamp.</p> <p><strong>Columns</strong></p> <p>The following columns are available:</p> <ul> <li>event: tag for query and language used for the query</li> <li>id: Tweet ID</li> <li>inserted_at: collection date</li> <li>last_updated_at: last update date (relevant for metrics such as follower count)</li> <li>text: Tweet text</li> <li>lang: language as determined by Twitter</li> <li>created_at: creation date of the Tweet</li> <li>conversation_id: Tweets with the same ID are part of the same reply tree to a tweet (provided by Twitter)</li> <li>author_follower_count: follower count of the Tweet's author account at the creation or last update time of the tweet</li> <li>replied_to: account the Tweet replies to</li> <li>replied_to_follower_count: follower count of the Tweet's replied to account at the creation or last update time of the tweet</li> <li>quoted: if quote tweet, ID of quoted tweet</li> <li>quoted_follower_count: analog to replied_to_follower_count</li> <li>retweeted: analog to quoted</li> <li>retweeted_follower_count: analog to replied_to_follower_count</li> <li>hashtags: hashtags of the Tweet</li> <li>urls: URLs in the tweet, shortened/unshortened, including links to media</li> <li>place_id: alphanumeric place ID provided by the Twitter API, mostly empty</li> </ul>

ShareScore

24/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
12
Harmonization
4
Access
0
Reuse readiness
0
Engagement
8

Topics