Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “political rights”
MIGR-TWIT Corpus. Migration Tweets of right and far-right politics in Europe
<p><strong>Description</strong></p> <p>The <strong>MIGR-TWIT Corpus</strong> is a multilingual corpus of tweets about the topic of migration in Europe. Within the framework of the collaborative research project OLiNDiNUM (Observatoire LINguistique du DIscours NUMérique, Linguistic Observatory of Online Debate) the MIGR-TWIT Corpus is created with the aim of developing language databases of online debate. Considering the global issue of migration in line with British and French political contexts of last dozen years from 2011 to 2022, the corpus consists of two sub-corpora: </p> <ul> <li> <p><strong>FR-R-MIGR-TWIT-2011-2022 Corpus </strong>for French language data (1 January 2011 - 30 June 2022) and </p> </li> <li> <p><strong>UK-R-MIGR-RA-TWIT-2012-2022 Corpus </strong>for English language data (1 January 2012 - 5 September 2022) <strong> </strong></p> </li> </ul> <p>Using the Twitter API v2 Academic Research, tweets containing at least one occurrence of migration or refugee related words are retrieved automatically from 28 right and far-right political figures and parties. The whole corpus contains 18,233 tweets and 533,198 words. </p> <p><strong>Scientific reference:</strong></p> <p>Pietrandrea, P., Battaglia, E. (2022). “Migrants and the EU”. The diachronic construction of ad hoc categories in French far-right discourse. Journal of Pragmatics 192, 139-157.</p> <p>Blandino, G. (2023). <em>10 years of public debate on immigration: combining topic modeling and corpus linguistics to examine the British (far-)right discourse on Twitter</em>, MA University of Wolverhampton</p> <p>Jeon, S. (2025). Le discours numérique sur l'immigration en France entre 2011 et 2022. Une analyse de corpus (Online Discourse on Immigration in France between 2011 and 2022. A Corpus Analysis), PhD Thesis, Université de Lille, France.</p> <p><strong>Contents</strong></p> <p>The whole corpus contains two CSV Zip files (tabular format) corresponding to each sub-corpus. The complete corpus is presented in two versions, one version with the tweet identifier (<strong><em>data__id</em></strong>) and the text of the tweet (<strong><em>data__text</em></strong>) as a header (folders named <em>FR-R-MIGR-TWIT-2011-2022_textonly</em> and <em>UK-R-MIGR-RA-TWIT-2012-2022_textonly</em>, respectively composed of 12 and 11 Zip files of every single year), and the other version with all tweet fields information included as a header, such as the posting date (<em><strong>data__created__at</strong></em>), the username (<strong><em>author__name</em></strong>), the number of retweets (<em><strong>data__public_metrics__retweet_count</strong></em>), etc., with two folders named <em>FR-R-MIGR-TWIT-2011-2022_meta</em> and <em>UK-R-MIGR-RA-TWIT-2012-2022_meta</em>. Detailed information for each sub-corpus is illustrated below.</p> <p><strong>1. FR-R-MIGR-TWIT-2011-2022 </strong></p> <ul> <li><strong>Created at: </strong>2022-08-08</li> <li> <p><strong>Language: </strong>FR<strong> </strong></p> </li> <li> <p><strong>Coverage: </strong>16 user accounts; 11,761 tweets; 358,491 words</p> </li> <li> <p><strong>Time of data collection: </strong>start=2011-01-01; end=2022-06-30 </p> </li> <li> <p><strong>Keywords: </strong>words derived from a latin root “<em><strong>migr</strong></em>” of <em>migrare</em></p> </li> <li> <p><strong>Corpus composition: </strong></p> </li> </ul> <table> <tbody> <tr> <th> </th> <th>Political figure/party</th> <th>Username</th> <th>Tweets</th> <th>Year concerned</th> </tr> <tr> <th>1</th> <td>Michel Barnier</td> <td>@MichelBarnier</td> <td>31</td> <td>2017-22</td> </tr> <tr> <th>2</th> <td>Valérie Pécresse</td> <td>@vpecresse</td> <td>81</td> <td>2017-22</td> </tr> <tr> <th>3</th> <td>Rassemblement National</td> <td>@RNational_off</td> <td>3,347</td> <td>2017-22</td> </tr> <tr> <th>4</th> <td>Nicolas Dupont-aignan</td> <td>@dupontaignan</td> <td>663</td> <td>2011-22</td> </tr> <tr> <th>5</th> <td>Éric Ciotti</td> <td>@ECiotti</td> <td>1,007</td> <td>2012-22</td> </tr> <tr> <th>6</th> <td>Christian Estrosi</td> <td>@cestrosi</td> <td>137</td> <td>2011-22</td> </tr> <tr> <th>7</th> <td>Marine Le Pen</td> <td>@MLP_officiel</td> <td>1,650</td> <td>2011-22</td> </tr> <tr> <th>8</th> <td>Valérie Boyer</td> <td>@valerieboyer13</td> <td>837</td> <td>2012-22</td> </tr> <tr> <th>9</th> <td>Florian Philippot</td> <td>@f_philippot</td> <td>485</td> <td>2012-22</td> </tr> <tr> <th>10</th> <td>Xavier Bertrand</td> <td>@xavierbertrand</td> <td>70</td> <td>2017-22</td> </tr> <tr> <th>11</th> <td>Marion Maréchal</td> <td>@MarionMarechal</td> <td>479</td> <td>2012-17,19-22</td> </tr> <tr> <th>12</th> <td>Philippe Meunier</td> <td>@Meunier_Ph</td> <td>245</td> <td>2013-22</td> </tr> <tr> <th>13</th> <td>Jordan Bardella</td> <td>@J_Bardella</td> <td>1,095</td> <td>2013-22</td> </tr> <tr> <th>14</th> <td>Nicolas Bay</td> <td>@NicolasBay_</td> <td>1,260</td> <td>2017-22</td> </tr> <tr> <th>15</th> <td>Emmanuel Macron</td> <td>@EmmanuelMacron</td> <td>72</td> <td>2017-22</td> </tr> <tr> <th>16</th> <td>Éric Zemmour</td> <td>@ZemmourEric</td> <td>302</td> <td>2019-22</td> </tr> <tr> <th>17</th> <td>Jean Messiha*</td> <td>Banned from Twitter (since July 2021)</td> <td>-</td> <td>-</td> </tr> </tbody> </table> <ul> <li>Political figures and parties of table above are listed in chronological order according to the dates on which they posted their first tweet.</li> <li> <p><strong>*</strong>Before the launching of Twitter API v2 Academic Research, migr-tweets were collected from the database of Europresse.com including 1,453 tweets of Jean Messiha as part of the reference study (Pietrandrea & Battaglia 2022). However, the Twitter account in question has been permanently banned since July 2021. For our data collection using the Twitter API started in September 2021, we could not access this account. Therefore, we decided not to include his tweets in the FR-R-MIGR-TWIT-2011-2022 for the sake of consistency with the rest of twitter data that are automatically retrieved.</p> </li> <li> <p>The sub-corpus FR-R-MIGR-TWIT-2017-2022 is developed, annotated and analyzed as part of a doctoral thesis in progress (<a href="https://theses.fr/s360032">Jeon, 2025</a>) with the aim of studying the semantic construction of migr-lexicon over the period between 2011 and 2022. </p> </li> </ul> <p><strong> </strong></p> <p><strong>2. UK-R-MIGR-RA-TWIT-2012-2022 </strong></p> <ul> <li> <p><strong>Created at: </strong>2022-09-06</p> </li> <li> <p><strong>Language: </strong>EN</p> </li> <li> <p><strong>Coverage: </strong>12 user accounts; 6,472 tweets; 174,707 words </p> </li> <li> <p><strong>Time of data collection: </strong>start=2012-01-01; end=2022-09-05</p> </li> <li> <p><strong>Keywords: </strong>words derived from a latin root “<strong><em>migr</em></strong>” of <em>migrare </em>in addition to the keywords “<strong><em>refugee</em></strong>(<strong><em>s</em></strong>)” and “<strong><em>asylum</em></strong>”.</p> </li> <li> <p><strong>Corpus composition:</strong></p> </li> </ul> <table> <tbody> <tr> <th> </th> <th>Political figure/party</th> <th>Username</th> <th>Tweets</th> <th>Year concerned</th> </tr> </tbody> <tbody> <tr> <th>1</th> <td>David Cameron</td> <td>@David_Cameron</td> <td>32</td> <td>2012-22</td> </tr> <tr> <th>2</th> <td>Amber Rudd</td> <td>@AmberRuddUK</td> <td>29</td> <td>2012-22</td> </tr> <tr> <th>3</th> <td>Sajid Javid</td> <td>@sajidjavid</td> <td>84</td> <td>2012-22</td> </tr> <tr> <th>4</th> <td>Boris johnson</td> <td>@BorisJohnson</td> <td>80</td> <td>2015-22</td> </tr> <tr> <th>5</th> <td>Priti Patel</td> <td>@pritipatel</td> <td>304</td> <td>2012-22</td> </tr> <tr> <th>6</th> <td>UK Home Office</td> <td>@ukhomeoffice</td> <td>909</td> <td>2012-22</td> </tr> <tr> <th>7</th> <td>Nigel Farage</td> <td>@Nigel_Farage</td> <td>1,010</td> <td>2012-22</td> </tr> <tr> <th>8</th> <td>Richard Tice</td> <td>@TiceRichard</td> <td>180</td> <td>2013-22</td> </tr> <tr> <th>9</th> <td>UKIP</td> <td>@UKIP</td> <td>2,746</td> <td>2012-22</td> </tr> <tr> <th>10</th> <td>Neil Hamilton</td> <td>@NeilUKIP</td> <td>252</td> <td>2013-22</td> </tr> <tr> <th>11</th> <td>Nick Griffin</td> <td>@NickGriffinBU</td> <td>542</td> <td>2012-22</td> </tr> <tr> <th>12</th> <td>Robin Tilbrook</td> <td>@RobinTilbrook</td> <td>304</td> <td>2012-22</td> </tr> </tbody> </table> <p> </p> <ul> <li> <p>2 out of 12 accounts are official accounts belonging to the” UK Home Office” department and the “UKIP” (United Kingdom Independence Party) party. 10 out of 12 accounts are political figures’ accounts.</p> </li> <li> <p>The corpus UK-R-MIGR-RA-TWIT-2012-2022 will be exploited for the following master’s thesis: Blandino, G. (2023). <em>10 years of public debate on immigration: combining topic modeling and corpus linguistics to examine the British (far-)right discourse on Twitter</em>, MA University of Wolverhampton.</p> </li> </ul> <p> </p>
Data for manuscript "The Prevalence of Terms Denoting Far-right and Far-left Political Extremism in U.S. and U.K. News Media"
<p>This data set belongs to an academic manuscript examining longitudinally (2000-2019) the prevalence of terms denoting far-right and far-left political extremism in a large corpus of more than 32 million written news and opinion articles from 54 news media outlets popular in the United States and the United Kingdom.</p> <p>The textual content of news and opinion articles from the 54 outlets listed in the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. This threshold was chosen to maximize inclusion in our analysis of outlets with sparse amounts of articles text per year. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript </p> <p>-articlesContainingTargetWords.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p> </p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions failed to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.</p> <p>Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word "Facebook" in an article footnote encouraging article readers to follow the journalist’s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. Some additional outlet-specific inaccuracies that we could identify occurred in "The Hill" and "Newsmax" news outlets where XPath expressions had some shortfalls at precisely capturing articles’ content. For "The Hill", in years 2007-2009, XPath expressions failed to capture the complete text of the article in about 40% of the articles. This does not necessarily result in incorrect frequency counts for that outlet but in a sample of articles’ words that is about 40% smaller than the total population of articles words for those three years. In the case of "NewsMax", the issue was that for some articles, XPath expressions captured the entire text of the article twice. Notice that this does not result in incorrect frequency counts. If a word appears x times in an article with a total of y words, the same frequency count will still be derived when our scripts count the word 2x times in the version of the article with a total of 2y words.</p> <p>To conclude, in a data analysis of 32 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 in the main manuscript for illustration of the accuracy of the frequency counts).</p>
Replication package for: Economic and Social Outsiders but Political Insiders: Sweden's Populist Radical Right
<p>This package contains all the code necessary to replicate the figures and tables in Dal Bo, E., F. Finan, O. Folke, T. Persson, and J. Rickne (forthcoming). "Economic and Social Outsiders but Political Insiders: Sweden's Populist Radical Right", Review of Economic Studies. Detailed instructions are also given about how to access the underlying data. </p>
Figure 1 from: Ek K, Goytia S, Lundmark C, Nysten-Haarala S, Pettersson M, Sandström A, Söderasp J, Stage J (2017) Challenges in Swedish hydropower – politics, economics and rights. Research Ideas and Outcomes 3: e21305. https://doi.org/10.3897/rio.3.e21305
Figure 1 - The challenges of two parallel water management systems.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.