Datasets from `Discovering and analysing lexical variation in social media text'
<p>This repository contains the datasets that were used in the following three papers, which are also included within P. Shoemark's PhD dissertation `Discovering and analysing lexical variation in social media text':</p> <ul> <li>P. Shoemark, D. Sur, L. Shrimpton, I. Murray, and S. Goldwater. <a href="https://www.aclweb.org/anthology/E17-1116/"><em>Aye or naw, whit dae ye hink? Scottish independence and linguistic identity on social media.</em></a> 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL). 2017. </li> <li>P. Shoemark, J. Kirby, and S. Goldwater. <a href="https://www.aclweb.org/anthology/W17-4908/"><em>Topic and audience effects on distinctively Scottish vocabulary usage in Twitter data.</em> </a>Workshop on Stylistic Variation at EMNLP. 2017.</li> <li>P. Shoemark, J. Kirby, and S. Goldwater. <a href="https://www.aclweb.org/anthology/W18-6101/"><em>Inducing a lexicon of sociolinguistic variables from code-mixed text.</em> </a>Workshop on Noisy User Generated Text at EMNLP. 2018. </li> </ul> <p>Datasets consist of tab-separated-values files, in which rows correspond to tweets, with columns for user ID, tweet ID, and timestamp. </p> <p>The text of the tweets (and associated metadata) can be re-downloaded (in batches of 100 per request) using Twitter's <a href="https://developer.twitter.com/en/docs/tweets/post-and-engage/api-reference/get-statuses-lookup">GET Statuses/Lookup</a> API endpoint <em>(NB: Tweets which have been deleted or made private since the original datasets were collected can <strong>not </strong>be re-downloaded, so it may not be possible to reconstruct the original datasets in their entirety). </em></p> <p> </p> <p>Most of these datasets were originally drawn from <a href="https://developer.twitter.com/en/docs/tweets/sample-realtime/api-reference/get-statuses-sample">the Sample endpoint of Twitter’s Streaming API (</a>a.k.a. the ‘Spritzer’), which provides a random 1% sample of all public tweets in near real-time:</p> <p> </p> <p><strong>GU Dataset: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-UK.zip?versionId=a9cc222e-1d07-4026-9b55-cf2189b58191">Geotagged-UK.zip</a></strong></p> <p><em>Tweets from Sept 2013 - Sept 2014 which are geotagged to locations within the UK.</em></p> <p>The file <strong>GU_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and are geotagged to locations within the UK.<strong> </strong><em><strong>• Tweets: </strong>1,768,334<strong> • Unique Users: </strong>455,075 <strong>• </strong></em></p> <p>The file <strong>GU.tsv </strong>contains the IDs for tweets in the <strong>final</strong> GU dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers. <em><strong>• Tweets: </strong>1,654,204<strong> • Unique Users: </strong>446,510 <strong>• </strong><sub>(the number of users in the GU dataset was slightly over-counted when reported in the paper; this is the actual number)</sub></em></p> <p> </p> <p><strong>GS Dataset: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-Scotland.zip?versionId=146a0589-ddc0-4c42-86f6-7da79686ae56">Geotagged-Scotland.zip </a></strong></p> <p><em>The subset of Tweets in the GU dataset which are geo-tagged to locations within Scotland, specifically.</em></p> <p>The file <strong>GS_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and are geotagged to locations within Scotland. <em><strong>• Tweets: </strong>178,401<strong> • Unique Users: </strong>41,685 <strong>• </strong></em></p> <p>The file <strong>GS.tsv </strong>contains the IDs for tweets in the <strong>final</strong> GS dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers. <em><strong>• Tweets: </strong>166,992<strong> • Unique Users: </strong>40,837 <strong>• </strong><sub>(the number of users in the GS dataset was slightly over-counted when reported in the paper; this is the actual number)</sub></em></p> <p> </p> <p><strong>IT Dataset & Controls: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Indyref-Tweets.zip?versionId=41042a3e-f247-4493-9abe-29fb9ae81ee1">Indyref-Tweets.zip</a></strong></p> <p><em>Tweets from Sept 2013 - Sept 2014 which contain hashtags relating to the 2014 Scottish Independence Referendum (plus 'control' tweets which are by the same users but do not contain referendum-related hashtags)</em></p> <p>The file <strong>IT_</strong><strong>pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and contain at least one of 47 hashtags we identified as relating to the 2014 Scottish Independence Referendum (see paper for hashtag list). <em><strong>• Tweets: </strong>77,708<strong> • Unique Users: </strong>26,019 <strong>• </strong></em></p> <p>The file <strong>IT</strong><strong>.tsv </strong>contains the IDs for tweets in the <strong>final</strong> IT dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, after applying additional pre-processing heuristics to filter out tweets by bots and spammers, and tweets which do not contain hashtags that we judged to <em>unambiguously</em> relate to the referendum. <em><strong>• Tweets: </strong>59,664 <strong>• Unique Users: </strong>18,589 <strong>• </strong></em></p> <p>The file <strong>IT_controls_pre-filtering.tsv </strong>contains the IDs for all tweets from the ‘Spritzer’ stream which were posted between September 1st 2013 and September 30th 2014, were classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets or quotes, and do <em><strong>not</strong></em> contain any of the hashtags we identified as relating to the 2014 Scottish Independence Referendum, but are authored by a user who has <em><strong>also</strong> </em>authored a tweet in <strong>IT_</strong><strong>pre-filtering.tsv</strong>. <em><strong>• Tweets: </strong>1,354,701 <strong>• Unique Users: </strong>26,019 <strong>• </strong></em></p> <p>The file <strong>IT_controls</strong><strong>.tsv </strong>contains the IDs for tweets in the <strong>final</strong> set of Control tweets that was used for the analyses in our <a href="http://www.aclweb.org/anthology/E17-1116/">EACL 2017</a> paper, i.e. tweets which do not contain referendum-related hashtags but are authored by users who also have also authored tweets in <strong>IT</strong><strong>.tsv</strong>. <em><strong>• Tweets: </strong>881,679 <strong>• Unique Users: </strong>18,589 <strong>• </strong></em></p> <p> </p> <p><strong>SG-Users’ and IH-Users' Autumn 2014 Timeline Datasets: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Autumn-2014_Timelines.zip?versionId=c456b530-361e-4b88-9006-5bebd4a43c92">Autumn-2014_Timelines.zip</a></strong></p> <p><em>Complete tweet histories from Aug-Oct 2014 for users from the GS and IT datasets.</em></p> <p>The file <strong>SG-Users_Autumn_2014_timelines_pre-filtering.tsv </strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the GS dataset, i.e. users we know to have used Scottish geotags. This dataset is not restricted to tweets which appear in the ‘Spritzer’ sample; instead it consists of complete User Timelines for the months concerned, retrieved using the <a href="https://developer.twitter.com/en/docs/tweets/timelines/api-reference/get-statuses-user_timeline">statuses/user timeline</a> endpoint of Twitter’s REST API in March 2017. Because there are limits on the number of tweets that can be retrieved using this endpoint, we were not able to retrieve complete Autumn 2014 tweet histories for <em>all</em> of the users in the GS dataset. <em><strong>• Tweets: </strong>3,014,029 </em> <em><strong>• Unique Users: </strong>18,274 <strong>• </strong></em></p> <p>The file <strong>SG-Users_Autumn_2014_timelines.tsv </strong>contains the IDs for tweets in the <strong>final</strong> SG-Users dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/W17-4908/">StyleVar 2017</a> paper. This dataset consists only of tweets which contain at least one instance of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details. <em><strong>• Tweets: </strong>1,112,931</em> <em><strong>• Unique Users: </strong>10,103 <strong>• </strong></em></p> <p>The file <strong>IH-Users_Autumn_2014_timelines_pre-filtering.tsv </strong>contains the IDs for tweets which were posted in August, September, or October 2014 by users from the IT dataset, i.e. users we know to have used Indyref-related hashtags. This dataset was collected in the same manner as SG-Users_Autumn_2014_timelines_pre-filtering.tsv; however, due to an error in this process, <strong>the IDs of most of the tweets in this dataset were not recorded</strong>. For such tweets the tweet ID column instead contains a placeholder tweet ID of the form _<user_ID>_<month>_<integer>, where the integer denotes the tweet's position in the reverse-chronological list of tweets that were retrieved for that user from that month (e.g. _147527441_09_286 is the placeholder tweet ID we assigned to the 286th September tweet we retrieved from the user whose Account ID is 147527441). Unfortunately, therefore, the tweets in this file whose 'IDs' begin with an underscore cannot be straightforwardly re-downloaded using Twitter's free <a href="https://developer.twitter.com/en/docs/tweets/post-and-engage/api-reference/get-statuses-lookup">GET Statuses/Lookup</a> API endpoint; but since their user IDs and timestamps are intact, it would still be possible to retrieve them using the (paid-for) <a href="http:// https://developer.twitter.com/en/docs/tutorials/choosing-historical-api">Historical APIs</a>. <em><strong>• Tweets: </strong>6,997,858 <strong>• Tweets whose IDs were recorded: </strong>288,394</em><em><strong> </strong> <strong>• Unique Users: </strong>14,645</em><em> <strong>• </strong></em></p> <p>The file <strong>IH-Users_Autumn_2014_timelines.tsv </strong>contains the IDs for tweets in the <strong>final</strong> IH-Users dataset that was used for the analyses in our <a href="http://www.aclweb.org/anthology/W17-4908/">StyleVar 2017</a> paper. Like the SG-Users dataset, this dataset consists only of tweets which contain at least one instance of one of 50 lexical variables that were the focus of the study, and has also undergone various other filtering steps; see the paper for full details. As with the pre-filtered version, most of the tweet IDs are unfortunately missing in this dataset. <em><strong>• Tweets: </strong>2,165,320 <strong>• Tweets whose IDs were recorded: </strong>115,366</em><em><strong> </strong> <strong>• Unique Users: </strong>10,784 <strong>• </strong></em></p> <p> </p> <p> </p> <p><strong>US Geotags: <a href="https://zenodo.org/api/files/caf09b6a-e57a-4a01-b2f0-ce5ab15583d1/Geotagged-USA.zip">Geotagged-USA.zip</a></strong></p> <p><em>Tweets from June 2013 - July 2016 which are geotagged to locations within the USA.</em></p> <p>The file <strong>GUSA.tsv</strong> contains all tweets from the ‘Spritzer’ sample which were posted between June 30th 2013 to July 1st 2016, are classified as English by <a href="https://github.com/saffsd/langid.py">langid.py</a>, are not retweets, do not contain urls or embedded media, are not by users with more than 1000 friends or followers, and are geotagged to locations within the USA. This dataset (along with the GU Dataset) was used in our <a href="http://www.aclweb.org/anthology/W18-6101/">WNUT 2018</a> paper. <em><strong>• Tweets: </strong></em> <em>8,375,573 </em> <em><strong>• Unique Users: </strong>1</em>,<em>826,260</em><em> <strong>• </strong></em></p> <p> </p>
ShareScore
20/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 12
- Reuse readiness
- 0
- Engagement
- 0