Monthly Samples of German Tweets (2019 - 2022)
<p><strong>Due to size limitations, this dataset is no longer updated. Further data as of January 2013 can be found in this dataset: <a href="https://doi.org/10.5281/zenodo.7670097">https://doi.org/10.5281/zenodo.7670097</a></strong></p> <p>This dataset contains German tweets and Twitter accounts recorded from the public Twitter Streaming API using the following filters:</p> <ul> <li>terms: <em>'a'</em>, <em>'e'</em>, <em>'i'</em>, <em>'o'</em>, <em>'u'</em>, and <em>'n'</em></li> <li>language: <em>'de'</em></li> </ul> <p>This filter combination should record a 1% sample of (almost) all German tweets (in German it is very unlikely that terms do not contain vowels or the frequently used character <em>'n'</em>).</p> <p>This dataset might be useful for the following use cases:</p> <ul> <li>Natural language processing (focussing on Twitter specifics in German, there exist only little German datasets)</li> <li>Social Network Analysis (Twitter network)</li> <li>Identifying behavioural patterns (retweeting, quoting, replying, hate speech, ...)</li> <li>Sharing political (or other domain-specific) content</li> <li>Bot detection</li> <li>and more ...</li> </ul> <p>This dataset will be updated monthly. Each sample (starting in April 2019) will follow the following naming pattern:</p> <ul> <li>german-tweet-sample-<em><YEAR></em>-<em><MONTH></em>.zip (size: ~ 1GB)</li> </ul> <p>It will contain several bunches of recorded JSON gzipped files.</p>
ShareScore
16/100
Overall dataset sharing score