Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
82
datasets available to search
ShareScore release 0.9.0
Dataset results
82 results for “newspapers”
Data from: Newspaper coverage of maternal health in Bangladesh, Rwanda, and South Africa: a quantitative and qualitative content analysis
Open the record for dataset details and reuse information.
Data from: Public opinion in Japanese newspaper readers’ posts under the prolonged COVID-19 infection spread 2019-2021: Contents analysis using Latent Dirichlet Allocation
Open the record for dataset details and reuse information.
Frequencies per million words for 5 epidemiologically relevant search terms in a dozen British 19th century newspapers
Open the record for dataset details and reuse information.
Data from: Exploring media representation of the exotic pet trade, with a focus on welfare: Taxonomic, framing, and language biases in peer-reviewed publications and newspaper articles
Open the record for dataset details and reuse information.
China Newspaper Directory
<p>Comprehensive directory of approximately 1,000 general-interest Chinese newspapers from 1981 to 2011.</p> <p>The newspaper directory is constructed from four data sources: (1) the Chinese Newspaper Directory (2003, 2006, 2010), published by the State Administration for Press and Publication (SPPA) -- the authority that issues newspaper licenses; (2) the Annual China Journalism Yearbooks (1982-2011), published by the Chinese Academy of Social Science; (3) the China Newspaper Industry Yearbooks (2004-2011), published by a Beijing-based research institute; and (4) an eight-volume collection of the front pages of major newspapers on the date of first publication.</p> <p>Description of the variables in the data set:</p> <p>1. Province, prefecture, and county are the location of the headquarter of a newspaper.<br> 2. newspaper_id: a numerical identifier for each newspaper<br> 3. license_id: a numerical identity for a newspaper’s license issued by the Chinese National Press and Publication Administration. Typically, when a newspaper changes its name, its license remains the same and so does the license_id variable.<br> 4. newspaper_ch: the Chinese name of a newspaper<br> 5. newspaper_en: the English translation of the Chinese name of a newspaper<br> 6. Previous_name: the previous name of a newspaper if the newspaper ever changes names<br> 7. Supervisor: the direct supervisory organization of a newspaper<br> 8. Supervisor_type: the type of the supervisor of a newspaper, classified into one of the following categories<br> - Party: the supervisor is a CCP committee<br> - Government agency: the supervisor is a government department<br> - Mass organization: the supervisor is a mass organization, such as a professional association<br> - Nonnewspaper media: the supervisor is a non-newspaper media such as a broadcaster<br> - Parent newspaper: the supervisor is a newspaper that owns the newspaper<br> - Party media group: the supervisor is a media group supervised by a CCP committee<br> - Soe: the supervisor is a state-owned enterprise<br> 9. Head_unit: the unit that manages the newspaper. Most of the time, the head unit of a newspaper is the same as its supervisor. In some cases, the supervisor (e.g., a CCP committee) may delegate power to an organization (e.g., a party newspaper) to manage a newspaper.<br> 10. Admin_rank: the administrative rank of a newspaper, classified as “central”, “province”, “capital city (of a province)”, “prefecture (non-capital-city prefecture” and “county.”<br> 11. Ownership: residual claimant of a newspaper<br> 12. Founded_date: the date when a newspaper is founded.<br> 13. Republish_date: the date when a newspaper is republished after a period of suspension.<br> 14. Termination_date: the date when a newspaper is terminated.<br> 15. Category_content: is the content type of a newspaper, classified as one of the following categories:<br> - Daily: a general party daily, i.e., the mouthpiece of a CCP committee<br> - Evening: a newspaper whose title contains the words “evening newspaper”, typically quasi-commercialized general-interest newspaper and published in the afternoon<br> - Metro: commercial general interest newspaper published in the morning<br> - Digest_general: a general interest digest, which does not publish original articles<br> - Mixed_econ: a semi-general interest newspaper with a focus on economic news<br> - Mixed_politics: a semi-general interest newspaper with a focus on political news<br> - Mixed_law: a semi-general interest newspaper with a focus on legal issues<br> - Mixed_life: a semi-general interest newspaper with a focus on the coverage of life style</p> <p> </p>
Digitized flood location dataset during cyclone Amphan from newspaper survey
<p>A dataset of inundation location, digitized and geotagged from newspaper reports during cyclone Amphan over Bengal delta.</p> <p>Description of columns:</p> <p>1. District: Location name at 2nd administrative level</p> <p>2. Upazila: Location name at 3rd administrative level</p> <p>3. Location: Location name</p> <p>4. Lon: Longitude</p> <p>5. Lat: Latitude</p> <p>6. Type: Type of flood - Flood or Inundation</p> <p>7. Mechanism: Type of flood/inundation mechanism. Breach (Embankment breaching), Hightide (Unembanked low-land), Overtopping (Embankment overflow)</p> <p>8. Source: Source newspaper name.</p> <p>9. Date: Corresponding date of the publishing of the news.</p> <p> </p> <p><strong>Acknowledgement:</strong></p> <p>CNES (through the TOSCA project BANDINO) and Embassy of France in Bangladesh for financial support. Also support from French research agency (Agence Nationale de la Recherche; ANR) under the DELTA project (ANR-17-CE03-0001).</p>
NEWSPAPER DISCOURSE AND ITS CHARACTERISTICS.
Open the record for dataset details and reuse information.
19th Century United States Newspaper Advert images with 'illustrated' or 'non illustrated' labels
<p>The Dataset contains images derived from the Newspaper Navigator (news-navigator.labs.loc.gov/), a dataset of images drawn from the Library of Congress Chronicling America collection (<a href="https://chroniclingamerica.loc.gov/">chroniclingamerica.loc.gov/</a>). </p> <blockquote> <p>[The Newspaper Navigator dataset] consists of extracted visual content for 16,358,041 historic newspaper pages in <em>Chronicling America</em>. The visual content was identified using an object detection model trained on annotations of World War 1-era Chronicling America pages, including annotations made by volunteers as part of the <a href="https://labs.loc.gov/work/experiments/beyond-words/">Beyond Words</a> crowdsourcing project.</p> <p>source:<a href="https://news-navigator.labs.loc.gov/"> https://news-navigator.labs.loc.gov/</a></p> </blockquote> <p>One of these categories is 'advertisements. This dataset contains a sample of these images with additional labels indicating if the advert is 'illustrated' or 'not illustrated'.</p> <p>The data is organised as follows:</p> <ul> <li>The images themselves can be found in `images.zip`</li> <li>`newspaper-navigator-sample-metadata.csv` contains metadata about each image drawn from the Newspaper Navigator Dataset.</li> <li>`ads.csv` contains the labels for the images as a CSV file</li> <li>`sample.csv` contains additional metadata about the images (based on the newspapers those images came from). </li> </ul> <p>This dataset was created for use in an under-review Programming Historian tutorial (<a href="http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt1">http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt1</a>) The primary aim of the data was to provide a realistic example dataset for teaching computer vision for working with digitised heritage material. The data is shared here since it may be useful for others. <strong>This data documentation is a work in progress and will be updated when the Programming Historian tutorial is released publicly. </strong></p> <p>The metadata CSV file contains the following columns:</p> <p>- filepath<br> - pub_date<br> - page_seq_num<br> - edition_seq_num<br> - batch<br> - lccn<br> - box<br> - score<br> - ocr<br> - place_of_publication<br> - geographic_coverage<br> - name<br> - publisher<br> - url<br> - page_url<br> - month<br> - year<br> - iiif_url</p>
Chinese newspaper coverage of human gene patents
<p>Workbook for the content analysis of Chinese newspaper coverage of human gene patents</p>
THE STATUS AND DISTINCTIVE FEATURES OF NEWSPAPER STYLE IN LANGUAGE
Open the record for dataset details and reuse information.
Data from: Women are seen more than heard in online newspapers
Feminist news media researchers have long contended that masculine news values shape journalists' quotidian decisions about what is newsworthy. As a result, it is argued, topics and issues traditionally regarded as primarily of interest and relevance to women are routinely marginalised in the news, while men's views and voices are given privileged space. When women do show up in the news, it is often as "eye candy," thus reinforcing women's value as sources of visual pleasure rather than residing in the content of their views. To date, evidence to support such claims has tended to be based on small-scale, manual analyses of news content. In this article, we report on findings from our large-scale, data-driven study of gender representation in online English language news media. We analysed both words and images so as to give a broader picture of how gender is represented in online news. The corpus of news content examined consists of 2,353,652 articles collected over a period of six months from more than 950 different news outlets. From this initial dataset, we extracted 2,171,239 references to named persons and 1,376,824 images resolving the gender of names and faces using automated computational methods. We found that males were represented more often than females in both images and text, but in proportions that changed across topics, news outlets and mode. Moreover, the proportion of females was consistently higher in images than in text, for virtually all topics and news outlets; women were more likely to be represented visually than they were mentioned as a news actor or source. Our large-scale, data-driven analysis offers important empirical evidence of macroscopic patterns in news content concerning the way men and women are represented.
Jump Entropy Models Trained on Ten Dutch Newspapers, 1950-1990
<p>Matrices of relative entropy (JSD) between front pages in ten Dutch newspapers published between 1950 and 1990. For more on the calculation of these time series see:</p> <p>MORE INFO ASAP</p> <p> </p>
UK newspaper data for McMenamin et al., Institutions and Elections, Journalism, 2021.
<p>UK newspaper data (CSV) for McMenamin et al., Institutions and Elections, Journalism, 2021.</p> <p>Code and other datasets for this article are also available on Zenodo.</p>
Data from: When scientific experts come to be media stars: an evolutionary model tested by analysing coronavirus media coverage across Italian newspapers
<p>This dataset includes metadata of the newspaper articles used for the paper "When scientific experts come to be media stars: an evolutionary model tested by analysing coronavirus media coverage across Italian newspapers". The dataset is in JSON format. The metadata includes: "uuid" (unique identifier we associated to an article), "URLs" (the URLs where the article was published), "sources" (newspaper and feed/section where the article was published), "datesPublished" (dates when the article was published/updated).</p> <p>License: Attribution-ShareAlike 4.0 International (<a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">https://creativecommons.org/licenses/by-sa/4.0/legalcode</a>)</p> <p> </p>
Excerpts from articles from the newspaper La Repubblica (2011).
<p><strong>Excerpts from articles from the newspaper <em>La Repubblica</em> (2011).</strong></p> <p>The search terms (dialett*, Italian*, lingu*) are highlighted.</p> <p>Intended for scientific use only. Copyright © La Repubblica and respective authors <a href="https://www.repubblica.it">https://www.repubblica.it</a></p>
Data from: Women are seen more than heard in online newspapers
Open the record for dataset details and reuse information.
NewsEye / READ AS training dataset from Austrian Newspapers (19th, early 20th C.)
<p>The dataset comprises Austrian newspaper pages from 19th and early 20th century with carefully annotated text. The page images were provided by the <a href="http://onb.ac.at/">Austrian National Library</a> and comprise 158 pages (training set). The data are formed according to the PAGE format (cf. Cf. <a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a> and the <a href="http://read.transkribus.eu/">READ </a>project. The guidelines with which the AS GT was created are uploaded here as well.</p>
Newspaper Rock
Utah Source: Objaverse 1.0 / Sketchfab
The Portrayal of Residential Care in Newspapers: Considerations for Resident Involvement
<p>Interview transcripts of reporters, administrators, and long-term care residents to understand the considerations for residents to be involved in the news media production process. </p> <p>Data contain sensitive health information and are therefore not made publicly available. </p>
Historic news comments from the G1 brazilian online newspaper
<p>This dataset contains data of user comments about the news published in the brazilian online newspaper G1 from 2010 and before. It contains 54,634 comments distributed between 4,549 news, which includes news published date, comments texts (in brazilian portuguese) and comments published date. That dataset intentionally ommits the titles, links, usernames or other kind of data that can redirect back to someones identity.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.