Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
82
datasets available to search
ShareScore release 0.9.0
Dataset results
82 results for “newspapers”
Training material - NBC- newspapers FR CH
<p>List of articles used to training a naive Bayes classifier, selected from French language Swiss newspapers, extracted from the impresso app. </p>
German newspaper data for McMenamin et al., Institutions and Elections, Journalism, 2021.
<p>German newspaper data (in STATA format) from McMenamin et al., Journalism, 2021.</p> <p>Code and other datasets also available on Zenodo.</p>
Forecasting with news sentiment: Evidence with UK newspapers
<p>These are datasets of economic sentiments derived from Uk newspapers using a dictionary and support vector machines. For more information on the application refer to : </p> <p>Rambaccussing, D. and Kwiatkowski, A., 2020. Forecasting with news sentiment: Evidence with UK newspapers. <em>International Journal of Forecasting</em>, <em>36</em>(4), pp.1501-1516.</p> <p>https://www.sciencedirect.com/science/article/pii/S0169207020300595</p> <p> </p> <p> </p>
Swedish Diachronic Corpus - Newspapers Part I
<p>The Swedish Diachronic Corpus (https://www2.lingfil.uu.se/person/pettersson/svediakorp/) is a project funded by Swe-Clarin (<a href="https://sweclarin.se/eng">https://sweclarin.se/eng</a>). The purpose of the project is to provide a corpus of texts covering the time period from Old Swedish to present day, with a wide variety of text types and freely available for download and search. The texts are provided in a plain text format and in a uniform CoNLL format, with placeholders for linguistic annotation. </p> <p>The dataset provided here is the first part of the newspaper section in the Swedish Diachronic Corpus. For other datasets within the corpus, see further the corpus website: https://www2.lingfil.uu.se/person/pettersson/svediakorp/. </p> <p>The project members are Eva Pettersson (Uppsala University) and Lars Borin (University of Gothenburg). For questions or comments, or if you are aware of any corpus resource that could be included in the Swedish Diachronic Corpus, don't hesitate to contact us!<br> <br> Eva Pettersson Department of Linguistics and Philology, Uppsala University eva.pettersson@lingfil.uu.se<br> Lars Borin Department of Swedish, University of Gothenburg lars.borin@svenska.gu.se</p>
Swedish Diachronic Corpus - Newspapers Part II
<p>The Swedish Diachronic Corpus (https://www2.lingfil.uu.se/person/pettersson/svediakorp/) is a project funded by Swe-Clarin (<a href="https://sweclarin.se/eng">https://sweclarin.se/eng</a>). The purpose of the project is to provide a corpus of texts covering the time period from Old Swedish to present day, with a wide variety of text types and freely available for download and search. The texts are provided in a plain text format and in a uniform CoNLL format, with placeholders for linguistic annotation. </p> <p>The dataset provided here is the second part of the newspaper section in the Swedish Diachronic Corpus. For other datasets within the corpus, see further the corpus website: https://www2.lingfil.uu.se/person/pettersson/svediakorp/. </p> <p>The project members are Eva Pettersson (Uppsala University) and Lars Borin (University of Gothenburg). For questions or comments, or if you are aware of any corpus resource that could be included in the Swedish Diachronic Corpus, don't hesitate to contact us!<br> <br> Eva Pettersson Department of Linguistics and Philology, Uppsala University eva.pettersson@lingfil.uu.se<br> Lars Borin Department of Swedish, University of Gothenburg lars.borin@svenska.gu.se</p>
Annotated files of the Soviet Belarusian newspaper Zviazda (years 1938, 1939, 1940)
<p>The present data contain the texts of the Soviet Belarusian newspaper 'Zviazda' for the years 1938, 1940 and 1940 as a ZIP archive.</p> <p>The data are in the <em>CoNLL-U</em> format, therefore morphologically and syntactically annotated.</p> <p>The data were generated from PDF files through OCR (by Tesseract) and anotated with UDPipe. These data were generated while exploring a pipeline to process Belarusian texts and they are very noisy.</p>
It’s not fur: newspaper article reporting of abandonment and relinquishment of pets exhibit taxonomic biases in framing and language use
Open the record for dataset details and reuse information.
Images from Newspaper Navigator predicted as maps, with human corrected labels
<p>The Dataset contains images derived from the Newspaper Navigator (news-navigator.labs.loc.gov/), a dataset of images drawn from the Library of Congress Chronicling America collection (chroniclingamerica.loc.gov/). </p> <blockquote> <p>[The Newspaper Navigator dataset] consists of extracted visual content for 16,358,041 historic newspaper pages in <em>Chronicling America</em>. The visual content was identified using an object detection model trained on annotations of World War 1-era Chronicling America pages, including annotations made by volunteers as part of the <a href="https://labs.loc.gov/work/experiments/beyond-words/">Beyond Words</a> crowdsourcing project.</p> <p>source:<a href="https://news-navigator.labs.loc.gov/"> https://news-navigator.labs.loc.gov/</a></p> </blockquote> <p>One of these categories is 'maps'. In the original training data for Newspaper Navigator, there were relatively few labelled examples of maps. The predictions for maps have an <a href="https://github.com/LibraryOfCongress/newspaper-navigator">Average Precision of 69.5%, and 34 images in the validation data</a>.</p> <p>This dataset contains a sample of these images which have been predicted as 'maps'. It also includes additional labels which indicate whether the predicted map image is a 'map' or 'not a map'. </p> <p>The data is organised as follows:</p> <ul> <li>The images themselves can be found in 'newspaper_maps.zip' </li> <li>`2020_30_10_13_19_228_sample.json` contains metadata about each image drawn from the Newspaper Navigator Dataset.</li> <li>map_labels.csv contains the labels for the images as a CSV file </li> </ul>
Data from: Newspaper coverage of maternal health in Bangladesh, Rwanda, and South Africa: a quantitative and qualitative content analysis
Objective: To examine newspaper coverage of maternal health in three countries that have made varying progress towards Millennium Development Goal 5 (MDG 5): Bangladesh (on track), Rwanda (making progress, but not on track) and South Africa (no progress). Design: We analysed each country's leading national English-language newspaper: Bangladesh's The Daily Star, Rwanda's The New Times/The Sunday Times, and South Africa's Sunday Times/The Times. We quantified the number of maternal health articles published from 1 January 2008 to 31 March 2013. We conducted a content analysis of subset of 190 articles published from 1 October 2010 to 31 March 2013. Results: Bangladesh's The Daily Star published 579 articles related to maternal health from 1 January 2008 to 31 March 2013, compared to 342 in Rwanda's The New Times/The Sunday Times and 253 in South Africa's Sunday Times/The Times over the same time period. The Daily Star had the highest proportion of stories advocating for or raising awareness of maternal health. Most maternal health articles in The Daily Star (83%) and The New Times/The Sunday Times (69%) used a 'human-rights' or 'policy-based' frame compared to 41% of articles from Sunday Times/The Times. Conclusions: In the three countries included in this study, which are on different trajectories towards MDG 5, there were differences in the frequency, tone and content of their newspaper coverage of maternal health. However, no causal conclusions can be drawn about this association between progress on MDG 5 and the amount and type of media coverage of maternal health.
Replication package for: "The Impact of Online Competition on Local Newspapers: Evidence from the Introduction of Craigslist"
<p>Djourelova, Milena, Durante, Ruben, and Gregory J. Martin (2023). "The Impact of Online Competition on Local Newspapers: Evidence from the Introduction of Craigslist."</p><p>The replication package contains software and datasets needed to produce tables and figures in the paper, as well as data cleaning scripts and raw data.</p><p>The replication package is also available as a Git repository at <a href="https://code.stanford.edu/gjmartin/craigslist-replication-code-and-data">https://code.stanford.edu/gjmartin/craigslist-replication-code-and-data</a></p>
Data for manuscript: "The Prevalence of Prejudice Denoting Terms in Spanish Newspapers"
<p>This data set contains frequency counts of target words in 5 million news and opinion articles from 3 popular newspapers in Spain: El País, El Mundo and ABC. The target words are listed in the associated manuscript and are mostly words that denote some type of prejudice. A few additional words not denoting prejudice are also available since they are used in the manuscript for illustration purposes.</p> <p>The textual content of news and opinion articles from the outlets listed in Figure 1 of the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We used derived word frequency counts from original sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript </p> <p>-targetWordsInArticlesCounts.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions can fail to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles might not be precise. </p> <p>To conclude, in a data analysis of millions of news articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 2 of main manuscript for supporting evidence).</p>
Frequencies per million words for 5 epidemiologically relevant search terms in a dozen British 19th century newspapers
<p>COVID-19 is the first known coronavirus pandemic. Nevertheless, the seasonal circulation of the four milder coronaviruses of humans – OC43, NL63, 229E and HKU1 – raises the possibility that these viruses are the descendants of more ancient coronavirus pandemics. This proposal arises by analogy to the observed descent of seasonal influenza subtypes H2N2 (now extinct), H3N2 and H1H1 from the pandemic strains of 1957, 1968 and 2009, respectively. Recent historical revisionist speculation has focussed on the influenza pandemic of 1889-1892, based on molecular phylogenetic reconstructions that show the emergence of human coronavirus OC43 around that time, probably by zoonosis from cattle. If the "Russian influenza", as The Times named it in early 1890, was not influenza but caused by a coronavirus, the origins of the other three milder human coronaviruses may also have left a residue of clinical evidence in the 19th century medical literature and popular press. In this paper, we search digitised 19th century British newspapers for evidence of previously unsuspected coronavirus pandemics. We conclude that there is little or no corpus linguistic signal in the UK national press for large-scale outbreaks of unidentified respiratory disease for the period 1785 to 1890.</p>
Y2K Newspaper Article
Y2K Newspaper Article Source: Objaverse 1.0 / Sketchfab
HOCON34k: A Corpus of Hate speech in Online Comments from German Newspapers
<p>We have compiled a dataset containing 34,223 comments in German, authored by users from online-platforms associated with public discourse in German newspapers. Each comment was annotated for hate speech and the adequacy of contextual information by a group of 29 volunteers, using a binary annotation approach. The inter-rater reliability for hate speech is 0.4428 across all annotators and increases to 0.6078 when considering an optimized subset of 12 annotators, as measured by Fleiss’ Kappa. Additionally, we present a baseline text classification using BERT, achieving an MCC-score up to 0.32 and an F2-score up to 0.64 in our initial experiment on this new corpus. The data set, named HOCON34k, comprising German hate speech comments from newspapers, is publicly available for research purposes.</p>
An Alien in the Newsroom: AI Anxiety in European and American Newspapers (dataset)
Open the record for dataset details and reuse information.
Handwritten newspapers used in Turunen, R. (2021). Shades of Red: Evolution of the Political Language of Finnish Socialism from the Nineteenth Century until the Civil War of 1918
<p>This dataset contains XLSX files of five different handwritten newspapers: <em>Kuritus </em>(1909–1911), <em>Palveliatar</em> (1907–1913, 1917), <em>Yritys</em> (1915–1917), <em>Tehtaalainen</em> (1908–1914, 1917) and <em>Nuija </em>(1899–1903, 1907–1909, 1912, 1914–1915).</p> <p>The original sources and the coding scheme for the XLSX files are described in:</p> <p>Turunen, R. (2021). <em>Shades of Red: Evolution of the Political Language of Finnish Socialism from the Nineteenth Century until the Civil War of 1918</em>. The Finnish Society for Labour History.</p>
Newspaper Coverage and Framing of Bats, and Their Impact on Readership Engagement
<p>Dataset underpinning the following study: </p> <p>Abstract: The media is a valuable pathway for transforming people’s attitudes towards conservation issues. Understanding how bats are framed in the media is hence essential for bat conservation, particularly considering the recent fearmongering and misinformation about the risks posed by bats. We reviewed bat-related articles published online no later than 2019 (before the recent COVID19 pandemic), in 15 newspapers from the five most populated countries in Western Europe. We examined the extent to which bats were presented as a threat to human health and the assumed general attitudes toward bats that such articles supported. We quantified press coverage on bat conservation values and evaluated whether the country and political stance had any information bias. Finally, we assessed their terminology, and, for the first time, modelled the active response from the readership based on the number of online comments. Out of 1096 articles sampled, 17% focused on bats and diseases, 53% on a range of ecological and conservation topics, and 30% only mention bats anecdotally. While most of the ecological articles did not present bats as a threat (97%), most articles focusing on diseases did so (80%). Ecosystem services were mentioned on very few occasions in both types (<30%), and references to the economic benefits they provide were meagre (<4%). Disease-related concepts were recurrent, and those articles that framed bats as a threat were the ones that garnered the highest number of comments. Therefore, we encourage the media to play a more proactive role in reinforcing positive conservation messaging by presenting the myriad ways in which bats contribute to safeguarding human well-being and ecosystem functioning.</p>
Data from: Public opinion in Japanese newspaper readers' posts under the prolonged COVID-19 infection spread 2019-2021: Contents analysis using Latent Dirichlet Allocation
<p><span>These data are based on free description from <span>readers' posts on Japanese hardcopy newspaper articles in the public domain</span>. </span><span>Upon searching for "coronavirus," we found 412 reader submissions during the April 7, 2020 to May 25, 2020, 521 during January 8, 2021 to March 21, 2021, 458 during the April 25, 2021 to June 6, 2021, and 519 during July 12, 2021 to September 30, 2021. </span><span>Topics of content were extracted by Latent Dirichlet Allocation, and then calculating the ratio of topic occurrence in each description.</span></p>
BL Newspapers sample plain-text data
<p>A dataset of .csv files each containing article texts from newspapers published on the Shared Research Repository. </p>
The Effect of Newspaper Reporting on COVID-19 Vaccine Hesitancy: a Randomised Controlled Trial
ClinicalTrials.gov study NCT05582564. IPD Sharing: NO. Countries: 1. Publications: 4.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.