Skip to main content
zenodorestricted

Monthly Samples of German Tweets (2019 - 2022)

<p><strong>Due to size limitations, this dataset is no longer updated. Further data as of January 2013 can be found in this dataset: <a href="https://doi.org/10.5281/zenodo.7670097">https://doi.org/10.5281/zenodo.7670097</a></strong></p> <p>This dataset contains German tweets and Twitter accounts recorded from the public Twitter Streaming API&nbsp;using the following filters:</p> <ul> <li>terms: <em>&#39;a&#39;</em>, <em>&#39;e&#39;</em>, <em>&#39;i&#39;</em>, <em>&#39;o&#39;</em>, <em>&#39;u&#39;</em>, and <em>&#39;n&#39;</em></li> <li>language: <em>&#39;de&#39;</em></li> </ul> <p>This filter combination should record a 1% sample of (almost) all German tweets (in German it is very unlikely that terms do not contain vowels or the frequently&nbsp;used character <em>&#39;n&#39;</em>).</p> <p>This dataset might be useful for the following use cases:</p> <ul> <li>Natural language processing (focussing on Twitter specifics in German, there exist only little German datasets)</li> <li>Social Network Analysis (Twitter network)</li> <li>Identifying behavioural patterns (retweeting, quoting, replying, hate speech, ...)</li> <li>Sharing political (or other domain-specific) content</li> <li>Bot detection</li> <li>and more ...</li> </ul> <p>This dataset will be updated monthly. Each sample (starting in April 2019) will follow the following naming pattern:</p> <ul> <li>german-tweet-sample-<em>&lt;YEAR&gt;</em>-<em>&lt;MONTH&gt;</em>.zip (size: ~ 1GB)</li> </ul> <p>It will contain several bunches of recorded JSON gzipped files.</p>

ShareScore

16/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
0
Reuse readiness
0
Engagement
4

Topics