Comprehensive Collection of English, German, Russian and Ukrainian Tweets Containing the Word or Hashtag Ukraine During the Russian Invasion, February 2022 until May 2023
<p>Comprehensive dataset of Tweets containing the keyword 'ukraine' (in German, Russian and Ukrainian) as well as '#ukraine' (in English) since the Russian Ukraine Invasion in February 2022. The user handle column has been excluded to protect deleted accounts that have not been retweeted or replied to. Tweets have been collected via the Academic API using the Search endpoint in four languages:</p> <table> <tbody> <tr> <td><strong>Language</strong></td> <td><strong>Query</strong></td> <td><strong> Name in Dataset (column: event) </strong></td> <td><strong>Number of Tweets<br></strong></td> </tr> <tr> <td>English</td> <td>#ukraine AND lang='en' </td> <td>ukraine-en-hashtag</td> <td>45.8 million</td> </tr> <tr> <td>German</td> <td>ukraine AND lang='de' </td> <td>ukraine</td> <td>19.6 million</td> </tr> <tr> <td>Russian</td> <td>Украина AND lang:ru </td> <td>ukraine-ru</td> <td>5.02 million</td> </tr> <tr> <td>Ukrainian</td> <td>Україна AND lang:uk </td> <td>ukraine-uk</td> <td>4.1 million</td> </tr> </tbody> </table> <p><strong>Collection dates</strong></p> <p>Details on collection dates per Tweet (e.g. to compare with creation dates) as well as the IDs of Tweets for consistency checks can be found here: <a href="https://github.com/Leibniz-HBI/ukraine_twitter_data">https://github.com/Leibniz-HBI/ukraine_twitter_data</a> (<a href="https://doi.org/10.17605/OSF.IO/RTQXN">https://doi.org/10.17605/OSF.IO/RTQXN</a>)</p> <p><strong>File Naming Scheme</strong></p> <p>To enable downloads of selected timeframes and languages, the files are named by language, start and end date of the tweet creation timestamp.</p> <p><strong>Columns</strong></p> <p>The following columns are available:</p> <ul> <li>event: tag for query and language used for the query</li> <li>id: Tweet ID</li> <li>inserted_at: collection date</li> <li>last_updated_at: last update date (relevant for metrics such as follower count)</li> <li>text: Tweet text</li> <li>lang: language as determined by Twitter</li> <li>created_at: creation date of the Tweet</li> <li>conversation_id: Tweets with the same ID are part of the same reply tree to a tweet (provided by Twitter)</li> <li>author_follower_count: follower count of the Tweet's author account at the creation or last update time of the tweet</li> <li>replied_to: account the Tweet replies to</li> <li>replied_to_follower_count: follower count of the Tweet's replied to account at the creation or last update time of the tweet</li> <li>quoted: if quote tweet, ID of quoted tweet</li> <li>quoted_follower_count: analog to replied_to_follower_count</li> <li>retweeted: analog to quoted</li> <li>retweeted_follower_count: analog to replied_to_follower_count</li> <li>hashtags: hashtags of the Tweet</li> <li>urls: URLs in the tweet, shortened/unshortened, including links to media</li> <li>place_id: alphanumeric place ID provided by the Twitter API, mostly empty</li> </ul>
ShareScore
24/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 12
- Harmonization
- 4
- Access
- 0
- Reuse readiness
- 0
- Engagement
- 8