Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

165

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

165 results for “politics”

Learn how ShareScore rates datasets ↗
zenodo40/100

CommonCrawl News Articles by Political Orientation

<p><strong>Dataset description &amp; reproduction steps</strong></p> <p>The dataset includes news articles gathered from CommonCrawl for media outlets that were selected based on their political orientation. The news articles span publication dates from 2010 to 2021. For more details, please check out our <a href="https://aclanthology.org/2022.findings-emnlp.152/">Paper</a> and <a href="https://github.com/webis-de/emnlp22-social-bias-representation-accuracy">GitHub repository</a>.</p> <p>The database file containing the news articles has two main tables, <em>article_urls</em> and <em>article_contents</em>. The tables have the following columns:</p> <p><em>article_urls</em>:</p> <ul> <li><code>uuid</code>: An ID that uniquely identifies this URL entry. This column is used as primary key for the table.</li> <li><code>url</code>: The plain text URL for the news article, as found in CommonCrawl.</li> <li><code>outlet_name</code>: The name of the news outlet that published the article.</li> </ul> <p><em>article_contents</em>:</p> <ul> <li><code>uuid</code>: An ID that uniquely identifies this content entry. This column is used as primary key for the table. The key is the same key used in the <em>article_urls</em> table to allow for cross-referencing.</li> <li><code>date</code>: The automatically extracted publishing date of the article. If it was not possible to automatically extract the date, this field remains empty.</li> <li><code>content</code>: The plain text content of the article automatically extracted from the crawled HTML document.</li> <li><code>content_preprocessed</code>: The articles content split by sentences.</li> <li><code>langauge</code>: The langauge of the article, as identified by the langdetect module (as ISO 639-1 code).</li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Assessing the consistency of fact-checking in political debates

<p>Dataset of the mixed-method study named &quot;Assessing the consistency of fact-checking in political debates&quot;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Quantifying polarization across political groups Twitter Dataset

<p>A CSV file consisting of Twitter data assessed between 2021 and 2022 from some of the members of the 116th US congress. The data has been used in studying political polarization on various domestic and geo-political policies.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Political, economic, and governance attitudes of blockchain users

<p>Survey responses for academic publication &quot;Political, economic, and governance attitudes of blockchain users&quot;</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Political map of the Arctic

<p>Political map of the Arctic -used as a source for the model of the Regional Security Complex in the Arctic</p> <p>The map was drafted as a part of research supported by Poland&rsquo;s National Centre for Science under the grant entitled: The adaptation of the regional security complex in the face of climate change: the example of the Arctic, with number: UMO-2019/35/N/HS5/00578.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Political, economic, and governance attitudes of blockchain users

<p>Survey data on the political, economic, and governance attitudes of blockchain users with accompanying description, visualization, and analysis.</p> <p>The&nbsp;<a href="https://metagov.typeform.com/cryptopolitics">Cryptopolitical Typology Quiz</a>&nbsp;was developed by the&nbsp;<a href="https://metagov.org/">Metagovernance Project</a>&nbsp;to help the crypto community understand its political, economic, and governance beliefs. Survey results were collected from September 27, 2021 through March 4, 2022 and have been published on the&nbsp;<a href="https://airtable.com/shr9LYMni8pBUVD6q/tblvwbt4KFm8MOSUQ">Govbase Airtable database</a>.</p> <p>This repository contains both a CSV export of the relevant results and the Python code used to visualize the distribution of responses, investigate the importance of blockchain affiliation and self-reported political orientation, and assess the validity of our constructed political score and types against any axes or features that emerge from the data.</p> <p>To view the results, check out the two Jupyter notebooks in this repository, paste the links to them into&nbsp;<a href="https://nbviewer.org/">nbviewer</a>&nbsp;for prettier in-browser viewing, or fork this repository and run them yourself!</p> <p>If you are interested in supporting ongoing work on the Cryptopolitics project, please get in touch with&nbsp;<a href="https://github.com/metagov/cryptopolitics-paper/blob/master/josh@metagov.org">josh@metagov.org</a>. To get involved with Metagov, join the Metagov&nbsp;<a href="https://metagov.pubpub.org/">community</a>&nbsp;or&nbsp;<a href="https://opencollective.com/metagov">staff</a>.</p> <p>The version in this release was used to generate the results and figures for a manuscript&nbsp;<a href="https://arxiv.org/abs/2301.02734">published on arXiv</a>&nbsp;and submitted for consideration for journal publication.</p> <p>Last modified on December 13, 2022.</p>

openmit-licenseFeb 2023View details →
zenodo40/100

Is plagiarism on the rise in Ethiopian politics? The case of three Master's theses at Addis Ababa University

<p>After an inquiry for plagiarism, the University of D&uuml;sseldorf in Germany <a href="https://www.theguardian.com/world/2013/feb/09/german-education-minister-quits-phd-plagiarism">revoked Annette Schavan&#39;s doctorate degree in 2013</a>. Schavan (CDU) served as Germany&#39;s Federal Minister of Education from 2005 to 2013, and resigned after her PhD was revoked. Karl-Theodor zu Guttenberg (CDU), Germany&#39;s defense minister, <a href="https://www.theguardian.com/world/2011/mar/01/german-defence-minister-resigns-plagiarism">resigned in 2011</a> when plagiarism was discovered in his PhD dissertation. As a result, the VroniPlag Wiki was established, where the level of plagiarism in German doctorate theses is investigated and documented via crowdsourcing. As a consequence, dozens of politicians were deprived of their degrees, voluntarily abandoned them, or fled politics. <a href="https://www.faz.net/aktuell/karriere-hochschule/hoersaal/franziska-giffey-verzicht-auf-doktortitel-ist-nicht-moeglich-18746134.html">Franziska Giffey (SPD) and Martin Huber (CSU)</a> are the most recent additions in 2023.</p> <p>The prestige of having an advanced degree is quite high in Ethiopia, as it is in Germany, and as a result, the temptation for shortcuts may be quite strong. We documented three instances of possible MSc thesis plagiarism at Addis Abeba University (AAU): Abraham Belay (2007), Takele Uma (2014), and Dagmawit Moges (2019), all of whom are (or were until recently) ministers in the Ethiopian government.</p> <p>The MSc theses were formally analysed for similarity with earlier works using <a href="https://www.turnitin.com/">Turnitin anti-plagiarism software</a>. The automated output was then further analysed. Similarities with works produced after or concurrently with the publishing date of the MSc theses were omitted from the similarity report; this pertains to a few instances of theses and articles in predatory journals that plagiarized passages from those theses. Also short similarities (eight words or less) were removed from the report. After that, the similarity scores remain high for the three MSc theses, with numerous fully copied paragraphs and sections.</p> <p><strong>Abraham Belay</strong> published an MSc thesis on &ldquo;<a href="http://etd.aau.edu.et/bitstream/handle/123456789/6598/Abraham%20belay.pdf?sequence=1&amp;isAllowed=y">DSP based vector control of induction motor</a>&rdquo; at AAU&rsquo;s Department of Electrical and Computer Engineering in 2007 (<a href="https://web.archive.org/web/20230124144908/http:/etd.aau.edu.et/bitstream/handle/123456789/6598/Abraham%20belay.pdf">archive</a>). In 2020, Abraham Belay (Prosperity Party) became minister of Innovation and Technology and in 2021 Minister of Defence. His MSc thesis presents 62% similarities with earlier work by other authors. See <strong>Annex A</strong>.</p> <p><strong>Takele Uma</strong> published an MSc thesis on &ldquo;<a href="http://etd.aau.edu.et/handle/123456789/11897">Environmental and Economic Benefits of Bioslurry from Coffee Husk Relative to Chemical Fertilizer</a>&rdquo; at AAU&rsquo;s School of Chemical and Bio-Engineering in 2014 (<a href="https://web.archive.org/web/20230409120611/http:/etd.aau.edu.et/bitstream/handle/123456789/11897/Takele%20Uma.pdf">archive</a>). Takele Uma (PP) was mayor of Addis Ababa from 2018 to 2020 and minister of Mines and Petroleum of Ethiopia from 2020 to 2023. His MSc thesis presents 38% similarities with earlier work. See <strong>Annex B</strong>.</p> <p><strong>Dagmawit Moges</strong> (PP) published an MSc thesis on &ldquo;<a href="http://etd.aau.edu.et/handle/123456789/19432">The Challenges and Prospects of Dynamic Electronic Service in Strengthening Demand-Responsive Transportation System in Addis Ababa, Ethiopia</a>&rdquo; at AAU&rsquo;s Department of Public Administration and Development Management in 2019 (<a href="https://web.archive.org/web/20230409120757/http:/etd.aau.edu.et/bitstream/handle/123456789/19432/Dagmawit%20Moges.pdf">archive</a>). Dagmawit was minister of Transport and Communications of Ethiopia from 2018 to 2023. In March 2023, she became director at the Peace Fund Secretariat of the African Union Commission. She <a href="https://smartermobility-africa.com/speakers/h-e-dagmawit-moges-bekele/">chairs also the boards</a> of Woldiya University, the Ethiopian Post, and the Ethiopian Roads Authority. Her MSc thesis presents 33% similarities with earlier work. <strong>See Annex C</strong>.</p> <p>We encourage readers to review the entire Turnitin reports for the three theses. It is not impossible that further paragraphs extracted from grey literature were missed by the Turnitin algorithms.</p> <p>On 28 February 2020, Addis Ababa University sent a <a href="https://archive.today/2023.04.08-084235/https:/t.me/Addisababauniversity/105">message in Amharic on its Telegram</a> channel, stating in substance that AAU has an anti-plagiarism policy, since plagiarism has been a significant hindrance to the University&#39;s efforts to provide excellent education. AAU further states that, earlier on, it has revoked a master&#39;s degree it had issued to a student due to plagiarism. It is only logical, then, that the three MSc theses annexed be formally scrutinized for plagiarism and the degrees possibly rescinded.</p> <p>Last but not least, examining the legitimacy of the country&#39;s politicians&#39; academic degrees would necessitate a collaborative effort, and we urge concerned academics, students, and alumni of Ethiopian universities to initiate a crowdsourcing effort to detect plagiarism in theses, as is done in Germany with <a href="https://en.wikipedia.org/wiki/VroniPlag_Wiki">VroniPlag Wiki</a> and in Russia with <a href="https://en.wikipedia.org/wiki/Dissernet">Dissernet</a>.</p> <p>&nbsp;</p> <p><strong>Acknowledgments</strong></p> <p>The authors acknowledge the owners of Twitter accounts @EgleGeek and @Zeate1 for pointing out that the three MSc theses addressed here reused texts from previous works. We would also like to thank Alex de Waal (Tufts University, Medford, MA, USA) and Boud Roukema (Nicolaus Copernicus University, Toruń, Poland) for exchanges of thoughts.</p> <p>&nbsp;</p> <p><strong>Annexes</strong></p> <p>Annex A. <a href="https://zenodo.org/record/7810624/files/Annex%20A%20-%20Abraham%20belay%20MSc%20thesis%20SimCheck2.pdf">Turnitin similarity report of the MSc thesis by Abraham Belay (2007)</a>.</p> <p>Annex B. <a href="https://zenodo.org/record/7810624/files/Annex%20B%20-%20Takele%20Uma%20MSc%20thesis%20similarity%20check.pdf">Turnitin similarity report of the MSc thesis by Takele Uma (2014).</a></p> <p>Annex C. <a href="https://zenodo.org/record/7810624/files/Annex%20C%20-%20Dagmawit%20Moges%20MSc%20thesis%20similarity%20check.pdf">Turnitin similarity report of the MSc thesis by Dagmawit Moges (2019)</a>.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Disgust and Politics: Investigating the Causal Mechanism Between Pathogen-Avoidance and Conservatism

<p>Data for the bachelor thesis of L. Y. Kogelheide. (.csv format)</p> <p>FEE_FEE are answers from the disgust sensitivity questionnaire</p> <p>Path are answers for the Disgust Image Set - Experimental Condition</p> <p>No_Path are answers for the Disgust Image Set - Control Condition</p> <p>Traditionalism_T are answers for the traditionalism questionnaire</p> <p>SDO_SDO are answers for the social dominance orientation questionnaire</p> <p>HC is honesty check</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Comments on YouTube videos of top two Indian political parties named Indian National Congress and Bhartiya Janata Party

<p>The datasets are taken from Various YouTube videos of top two Indian political parties named Indian National Congress and Bhartiya Janata Party.</p> <p>Both the datasets are divided into two categories: -</p> <p>Label 1- Positive</p> <p>Label 2- Negative</p> <p>All the labelling has been done manually.</p> <p><br> &nbsp;</p> <p><strong>Indian National Congress dataset:</strong></p> <p>Characteristics: Bivariate</p> <p>Number of instances in dataset: 1998</p> <p>Area of subject: Politics</p> <p>Attribute characteristics of dataset: Real</p> <p>Number of attributes: 2</p> <p>Date donated: March, 2019</p> <p>Task associated: Classification(binary)</p> <p>Missing values: Null</p> <p><br> &nbsp;</p> <p><strong>Bhartiya Janata Party dataset:</strong></p> <p>Characteristics: Bivariate</p> <p>Number of instances in dataset: 1952</p> <p>Area of subject: Politics</p> <p>Attribute characteristics of dataset: Real</p> <p>Number of attributes: 2</p> <p>Date donated: March, 2019</p> <p>Task associated: Classification(binary)</p> <p>Missing values: Null</p> <p><br> &nbsp;</p> <p><br> &nbsp;</p> <p><strong>Bothe datasets contains equal number of positive and negative comments:</strong></p> <p><br> &nbsp;</p> <p>Total number of positive comments present in Bhartiya Janata Party dataset =976</p> <p>Total number of negative comments present in Bhartiya Janata Party dataset =976</p> <p>Total number of positive comments present in Indian National Congress dataset=999</p> <p>Total number of negative comments present in Bhartiya Janata Party dataset=999</p> <p><br> &nbsp;</p> <p><strong>Both datasets contain following attributes:</strong></p> <ul> <li> <p><strong>comment text </strong></p> </li> <li> <p><strong>Labels </strong></p> </li> </ul>

opencc-by-4.0May 2019View details →
dryad40/100

Has the Supreme Court become just another political branch? Public perceptions of Court approval and legitimacy in a post-Dobbs world

Open the record for dataset details and reuse information.

publicJan 2024View details →
zenodo36/100

Money and Politics

<p>The data sets released here has been used in our study on &quot;The Geography of Money and Politics.&quot;</p> <p>In this study, we examined the social antecedents for contributing to campaigns, with a particular focus on the role of population density and social networking opportunities. Using ten years of US campaign contribution data from the Federal Election Commission (FEC) and a national survey of party leaders, we reported interesting findings regarding the interplay among density and mobility (operationalized by commuting flows). This analysis also reveals differences between political parties. Democrats are more dependent on social networking in dense population areas. This difference in the importance of social networking opportunities present in geographical space helps explain macro-level patterns in party fundraising.</p> <p>Our study was based on a collection of data sets, including: 1) FEC contribution, 2) US census, 3) US presidential vote share, 4) earning, and 5) commuting.&nbsp;</p> <p>These data were used to create the final analysis dataset consisting of variables in the models described in our R&amp;P paper.</p> <p>The MATLAB code can be used to reproduce the results w.r.t. all models specified in the paper.&nbsp;</p> <p>More details can be found in the enclosed README files.</p> <p>&nbsp;</p> <p><strong>Publication</strong></p> <p>If you make use of this data set, please cite:</p> <p>Lin, Y.-R., Kennedy, R., Lazer, D. (2017). The Geography of Money and Politics: Population Density, Social Networking and Political Contributions. Research &amp; Politics, 4(4) (doi: 10.1177/2053168017742015)</p>

openother-openJul 2017View details →
zenodo36/100

« Lend your Money, Lose your Friend? » - Chinese Official Lending and Bilateral Political Alignment: The Case of Africa

<p>This database provides a set of 45 variables related to UNGA voting affinity vis-à-vis China, official lending, and other bilateral economic and political indicators for China and 43 African countries over the period 2000-2020. The indicators are grouped into four categories: voting data, loan and debt, economic indicators, and political indicators.&nbsp;</p><p>This dataset was compiled in order to conduct research and econometric work for a journal article entitled:&nbsp;</p><p><strong>« Lend your Money, Lose your Friend? » - Chinese Official Lending and Bilateral Political Alignment: The Case of Africa</strong></p><p>, written by Clément Durif, Junior Resarch Fellow at the Asia Centre <a href="mailto:clement.durif@sciencespo.fr">clement.durif@sciencespo.fr</a><br><br>The status of this journal article is pending submittal and acceptation from a Journal Publication</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Analyzing Linguistic Patterns in the Social Media Discourse of Juan Guaidó and Nicolás Maduro during the 2019 Political Conflict in Venezuela - research data

<p>In this paper, the political conflict in Venezuela in 2019 is approached from a corpus linguistic point of view. The conflict between Juan Guaid&oacute;, the speaker of the parliament, and Nicol&aacute;s Maduro, who won the internationally unrecognized 2018 presidential election, escalated on January 23, 2019, when Guaid&oacute; proclaimed himself the legitimate president of Venezuela. By comparatively analyzing the tweets of the two politicians three months before and after January 23, 2019, a corpus-based discourse analysis will be conducted to investigate whether and to what extent linguistic patterns (especially most frequent words and their co-occurrences, as well as n-grams) change within the respective social media communication of these political opponents. The analysis reveals changes in linguistic patterns, especially with respect to co-occurrences and n-grams, detected in the corpus data, and demonstrates that politicians use Twitter to present themselves, in the case of Guaid&oacute;, as the representative of the people wanting to lead Venezuela into a democratic future, and, in the case of Maduro, as the only legitimate president and defender of Venezuela against internal and external threats.</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

Women's political participation in city council, Debretabor City Administration, Ethiopia: An exploration of barriers

<p><span>Women have traditionally been marginalized within the political sphere, notwithstanding their significant achievements. The study looks at the barriers that stand in the way of women participating actively in politics, with a particular emphasis on the Debre Tabor City Council. A case study research design was used in the study. The sample techniques used were maximum variation, deviant case, and intensity sampling. Semi-structured interviews, life histories, and focus group discussions (FGD) were also used to obtain the data. The findings showed that women's political engagement in city council was in danger due to a variety of factors, including <span>gender bias in decision-making processes</span>, the country's shifting political landscape, and <span>Islamic religion fellows' perspectives.</span> <span>In light of this, the research findings urged the need for gender-sensitive policy formulation within the city administration council. </span></span></p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Data for "Book bans in political context: Evidence from U.S. schools"

<p>This repository contains the data for the paper "Book bans in political context: Evidence from U.S. schools". All authors contributed equally and are listed alphabetically.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Toxic Sentence Classification Dataset with labels of categories such as religion, mental health, race, sex, body image, disability, physical abuse, and politics

<p>The dataset has a collection of various toxic sentences belonging to different categories. It was collected from various sources. It indicates which category each sentence belongs to. The values of the category columns are binary 1 or 0 indicating whether the sentence belongs to that particular category or not. Each sentence belongs to only 1 category.&nbsp;</p> <p>&nbsp;</p> <p>Columns:<br>1.comment_text: Contains toxic sentences that are insensitive and offensive, focusing on various categories.<br>2.mental_health: Binary value 1 indicates that the sentence focuses on mental health.<br>3.Race:Binary value 1 indicates that the sentence is racist.<br>4.sex:Binary value 1 indicates that the sentence focuses on sexuality.<br>5.body_image:Binary value 1 indicates that the sentence focuses on body image.<br>6.disability:Binary value 1 indicates that the sentence focuses on physical disability and related issues.<br>7.religion:Binary value 1 indicates that the sentence can be triggering to people who are extremely religious.<br>8.physical_abuse:Binary value 1 indicates that the sentence focuses on physical abuse issues.<br>9.politics:Binary value 1 indicates that the sentence focuses on political issues.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Documents used in the PLANET4B analysis of biodiversity discourse by political parties

<p>These files include press releases that have been published on the internet by European political parties, and which were used in the PLANET4B project analysis of the discourse on biodiversity.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

American Policy Conflict in the Hothouse: Exploring the Politics of Climate Inaction and Polycentric Rebellion

<p>Supporting data to Figure 6 in the forthcoming publication &quot;American Policy Conflict in the Hothouse: Exploring the Politics of Climate Inaction and Polycentric Rebellion&quot; in Energy Research &amp; Social Science (ERSS). Dataset provides insight into the components used to create Figure 6 and assess state-level policy contributions to carbon dioxide emission reductions (2002-2030).</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Data in Support of: Coal Exit Policy Must Confront Loopholes and Laggards for Political Momentum to Matter for Paris Targets

<p>Input and output data of COALogit and REMIND, which generated the dynamic feasibility space of national accession to the Powering Past Coal Alliance (PPCA) and the long-term energy system and emissions impacts of PPCA-induced coal phase-out scenarios, respectively.&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Data for manuscript "The Prevalence of Terms Denoting Far-right and Far-left Political Extremism in U.S. and U.K. News Media"

<p>This data set belongs to an academic manuscript examining longitudinally (2000-2019) the prevalence of terms denoting far-right and far-left political extremism in a large corpus of more than 32 million written news and opinion articles from 54 news media outlets popular in the United States and the United Kingdom.</p> <p>The textual content of news and opinion articles from the 54 outlets listed in the main manuscript is available in the outlet&#39;s online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions.&nbsp;Tokens were lowercased prior to estimating frequency counts.&nbsp;To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. This threshold was chosen to maximize inclusion in our analysis of outlets with sparse amounts of articles text per year.&nbsp;</p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript&nbsp;</p> <p>-articlesContainingTargetWords.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>&nbsp;</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions failed to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.</p> <p>Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word &quot;Facebook&quot; in an article footnote encouraging article readers to follow the journalist&rsquo;s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. Some additional outlet-specific inaccuracies that we could identify occurred in &quot;The Hill&quot; and &quot;Newsmax&quot; news outlets where XPath expressions had some shortfalls at precisely capturing articles&rsquo; content. For &quot;The Hill&quot;, in years 2007-2009, XPath expressions failed to capture the complete text of the article in about 40% of the articles. This does not necessarily result in incorrect frequency counts for that outlet but in a sample of articles&rsquo; words that is about 40% smaller than the total population of articles words for those three years. In the case of &quot;NewsMax&quot;, the issue was that for some articles, XPath expressions captured the entire text of the article twice. Notice that this does not result in incorrect frequency counts. If a word appears x times in an article with a total of y words, the same frequency count will still be derived when our scripts count the word 2x times in the version of the article with a total of 2y words.</p> <p>To conclude, in a data analysis of 32 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles&rsquo; content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 in the main manuscript for illustration of the accuracy of the frequency counts).</p>

opencc-by-4.0Sep 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record