Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

82

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

82 results for “newspapers”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Newspaper coverage of maternal health in Bangladesh, Rwanda, and South Africa: a quantitative and qualitative content analysis

Open the record for dataset details and reuse information.

publicDec 2015View details →
dryad32/100

Data from: Public opinion in Japanese newspaper readers’ posts under the prolonged COVID-19 infection spread 2019-2021: Contents analysis using Latent Dirichlet Allocation

Open the record for dataset details and reuse information.

publicAug 2023View details →
dryad32/100

Frequencies per million words for 5 epidemiologically relevant search terms in a dozen British 19th century newspapers

Open the record for dataset details and reuse information.

publicAug 2022View details →
dryad32/100

Data from: Exploring media representation of the exotic pet trade, with a focus on welfare: Taxonomic, framing, and language biases in peer-reviewed publications and newspaper articles

Open the record for dataset details and reuse information.

publicJun 2025View details →
zenodo28/100

China Newspaper Directory

<p>Comprehensive directory of approximately 1,000 general-interest Chinese newspapers from 1981 to 2011.</p> <p>The newspaper directory is constructed from four data sources: (1) the Chinese Newspaper Directory (2003, 2006, 2010), published by the State Administration for Press and Publication (SPPA) -- the authority that issues newspaper licenses; (2) the Annual China Journalism Yearbooks (1982-2011), published by the Chinese Academy of Social Science; (3) the China Newspaper Industry Yearbooks (2004-2011), published by a Beijing-based research institute; and (4) an eight-volume collection of the front pages of major newspapers on the date of first publication.</p> <p>Description of the variables in the data set:</p> <p>1.&nbsp;&nbsp; &nbsp;Province, prefecture, and county are the location of the headquarter of a newspaper.<br> 2.&nbsp;&nbsp; &nbsp;newspaper_id: a numerical identifier for each newspaper<br> 3.&nbsp;&nbsp; &nbsp;license_id: a numerical identity for a newspaper&rsquo;s license issued by the Chinese National Press and Publication Administration. Typically, when a newspaper changes its name, its license remains the same and so does the license_id variable.<br> 4.&nbsp;&nbsp; &nbsp;newspaper_ch: the Chinese name of a newspaper<br> 5.&nbsp;&nbsp; &nbsp; newspaper_en: the English translation of the Chinese name of a newspaper<br> 6.&nbsp;&nbsp; &nbsp;Previous_name: the previous name of a newspaper if the newspaper ever changes names<br> 7.&nbsp;&nbsp; &nbsp;Supervisor: the direct supervisory organization of a newspaper<br> 8.&nbsp;&nbsp; &nbsp;Supervisor_type: the type of the supervisor of a newspaper, classified into one of the following categories<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Party: the supervisor is a CCP committee<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Government agency: the supervisor is a government department<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Mass organization: the supervisor is a mass organization, such as a professional association<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Nonnewspaper media: the supervisor is a non-newspaper media such as a broadcaster<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Parent newspaper: the supervisor is a newspaper that owns the newspaper<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Party media group: the supervisor is a media group supervised by a CCP committee<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Soe: the supervisor is a state-owned enterprise<br> 9.&nbsp;&nbsp; &nbsp;Head_unit: the unit that manages the newspaper. Most of the time, the head unit of a newspaper is the same as its supervisor. In some cases, the supervisor (e.g., a CCP committee) may delegate power to an organization (e.g., a party newspaper) to manage a newspaper.<br> 10.&nbsp;&nbsp; &nbsp;Admin_rank: the administrative rank of a newspaper, classified as &ldquo;central&rdquo;, &ldquo;province&rdquo;, &ldquo;capital city (of a province)&rdquo;, &ldquo;prefecture (non-capital-city prefecture&rdquo; and &ldquo;county.&rdquo;<br> 11.&nbsp;&nbsp; &nbsp;Ownership: residual claimant of a newspaper<br> 12.&nbsp;&nbsp; &nbsp; Founded_date: the date when a newspaper is founded.<br> 13.&nbsp;&nbsp; &nbsp;Republish_date: the date when a newspaper is republished after a period of suspension.<br> 14.&nbsp;&nbsp; &nbsp;Termination_date: the date when a newspaper is terminated.<br> 15.&nbsp;&nbsp; &nbsp;Category_content: is the content type of a newspaper, classified as one of the following categories:<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Daily: a general party daily, i.e., the mouthpiece of a CCP committee<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Evening: a newspaper whose title contains the words &ldquo;evening newspaper&rdquo;, typically quasi-commercialized general-interest newspaper and published in the afternoon<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Metro: commercial general interest newspaper published in the morning<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Digest_general: a general interest digest, which does not publish original articles<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Mixed_econ: a semi-general interest newspaper with a focus on economic news<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Mixed_politics: a semi-general interest newspaper with a focus on political news<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Mixed_law: a semi-general interest newspaper with a focus on legal issues<br> &nbsp;&nbsp; &nbsp;-&nbsp;&nbsp; &nbsp;Mixed_life: a semi-general interest newspaper with a focus on the coverage of life style</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo28/100

Digitized flood location dataset during cyclone Amphan from newspaper survey

<p>A dataset of inundation location, digitized and geotagged from newspaper reports during cyclone Amphan over Bengal delta.</p> <p>Description of columns:</p> <p>1. District: Location name at 2nd administrative level</p> <p>2. Upazila: Location name at 3rd administrative level</p> <p>3. Location: Location name</p> <p>4. Lon: Longitude</p> <p>5. Lat: Latitude</p> <p>6. Type: Type of flood - Flood or Inundation</p> <p>7. Mechanism: Type of flood/inundation mechanism. Breach (Embankment breaching), Hightide (Unembanked low-land), Overtopping (Embankment overflow)</p> <p>8. Source: Source newspaper name.</p> <p>9. Date: Corresponding date of the publishing of the news.</p> <p>&nbsp;</p> <p><strong>Acknowledgement:</strong></p> <p>CNES (through the TOSCA project BANDINO) and Embassy of France in Bangladesh for financial support. Also support from French research agency (Agence Nationale de la Recherche; ANR) under the DELTA project (ANR-17-CE03-0001).</p>

opencc-by-4.0Oct 2020View details →
zenodo28/100

NEWSPAPER DISCOURSE AND ITS CHARACTERISTICS.

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

19th Century United States Newspaper Advert images with 'illustrated' or 'non illustrated' labels

<p>The Dataset contains images derived from the Newspaper Navigator (news-navigator.labs.loc.gov/), a dataset of images drawn from the Library of Congress Chronicling America collection (<a href="https://chroniclingamerica.loc.gov/">chroniclingamerica.loc.gov/</a>).&nbsp;</p> <blockquote> <p>[The Newspaper Navigator dataset] consists of extracted visual content for 16,358,041 historic newspaper pages in&nbsp;<em>Chronicling America</em>. The visual content was identified using an object detection model trained on annotations of World War 1-era Chronicling America pages, including annotations made by volunteers as part of the&nbsp;<a href="https://labs.loc.gov/work/experiments/beyond-words/">Beyond Words</a>&nbsp;crowdsourcing project.</p> <p>source:<a href="https://news-navigator.labs.loc.gov/"> https://news-navigator.labs.loc.gov/</a></p> </blockquote> <p>One of these categories is &#39;advertisements. This dataset contains a sample of these images with additional labels indicating if the advert is &#39;illustrated&#39; or &#39;not illustrated&#39;.</p> <p>The data is organised as follows:</p> <ul> <li>The images themselves can be found in `images.zip`</li> <li>`newspaper-navigator-sample-metadata.csv` contains metadata about each image drawn from the Newspaper Navigator Dataset.</li> <li>`ads.csv` contains the labels for the images as a CSV file</li> <li>`sample.csv` contains additional metadata about the images (based on the newspapers those images came from).&nbsp;</li> </ul> <p>This dataset was created for use in an under-review Programming Historian tutorial (<a href="http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt1">http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt1</a>) The primary aim of the data was to provide a realistic example dataset for teaching computer vision for working with digitised heritage material. The data is shared here since it may be useful for others. <strong>This data documentation is a work in progress and will be updated when the Programming Historian tutorial is released publicly. </strong></p> <p>The metadata CSV file contains the following columns:</p> <p>- filepath<br> - pub_date<br> - page_seq_num<br> - edition_seq_num<br> - batch<br> - lccn<br> - box<br> - score<br> - ocr<br> - place_of_publication<br> - geographic_coverage<br> - name<br> - publisher<br> - url<br> - page_url<br> - month<br> - year<br> - iiif_url</p>

openother-pdOct 2021View details →
zenodo28/100

Chinese newspaper coverage of human gene patents

<p>Workbook for the content analysis of Chinese newspaper coverage of human gene patents</p>

opencc-by-4.0Apr 2018View details →
zenodo28/100

THE STATUS AND DISTINCTIVE FEATURES OF NEWSPAPER STYLE IN LANGUAGE

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
dryad28/100

Data from: Women are seen more than heard in online newspapers

Feminist news media researchers have long contended that masculine news values shape journalists' quotidian decisions about what is newsworthy. As a result, it is argued, topics and issues traditionally regarded as primarily of interest and relevance to women are routinely marginalised in the news, while men's views and voices are given privileged space. When women do show up in the news, it is often as "eye candy," thus reinforcing women's value as sources of visual pleasure rather than residing in the content of their views. To date, evidence to support such claims has tended to be based on small-scale, manual analyses of news content. In this article, we report on findings from our large-scale, data-driven study of gender representation in online English language news media. We analysed both words and images so as to give a broader picture of how gender is represented in online news. The corpus of news content examined consists of 2,353,652 articles collected over a period of six months from more than 950 different news outlets. From this initial dataset, we extracted 2,171,239 references to named persons and 1,376,824 images resolving the gender of names and faces using automated computational methods. We found that males were represented more often than females in both images and text, but in proportions that changed across topics, news outlets and mode. Moreover, the proportion of females was consistently higher in images than in text, for virtually all topics and news outlets; women were more likely to be represented visually than they were mentioned as a news actor or source. Our large-scale, data-driven analysis offers important empirical evidence of macroscopic patterns in news content concerning the way men and women are represented.

opencc-zeroDec 2015View details →
zenodo28/100

Jump Entropy Models Trained on Ten Dutch Newspapers, 1950-1990

<p>Matrices of relative entropy (JSD) between front pages in ten Dutch newspapers published between 1950 and 1990. For more on the calculation of these time series see:</p> <p>MORE INFO ASAP</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo28/100

UK newspaper data for McMenamin et al., Institutions and Elections, Journalism, 2021.

<p>UK newspaper data (CSV) for McMenamin et al., Institutions and Elections, Journalism, 2021.</p> <p>Code and other datasets for this article are also available on Zenodo.</p>

opencc-by-4.0Oct 2021View details →
zenodo28/100

Data from: When scientific experts come to be media stars: an evolutionary model tested by analysing coronavirus media coverage across Italian newspapers

<p>This dataset includes metadata of the newspaper articles used for the paper &quot;When scientific experts come to be media stars: an evolutionary model tested by analysing coronavirus media coverage across Italian newspapers&quot;. The dataset is in JSON format. The metadata includes: &quot;uuid&quot; (unique identifier we associated to an article), &quot;URLs&quot; (the URLs where the article was published), &quot;sources&quot; (newspaper and feed/section where the article was published), &quot;datesPublished&quot; (dates when the article was published/updated).</p> <p>License: Attribution-ShareAlike 4.0 International (<a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">https://creativecommons.org/licenses/by-sa/4.0/legalcode</a>)</p> <p>&nbsp;</p>

openother-atMar 2023View details →
zenodo28/100

Excerpts from articles from the newspaper La Repubblica (2011).

<p><strong>Excerpts from articles from the newspaper <em>La Repubblica</em> (2011).</strong></p> <p>The search terms (dialett*, Italian*, lingu*) are highlighted.</p> <p>Intended for scientific use only. Copyright &copy; La Repubblica and respective authors&nbsp;<a href="https://www.repubblica.it">https://www.repubblica.it</a></p>

opencc-by-4.0Jul 2023View details →
dryad28/100

Data from: Women are seen more than heard in online newspapers

Open the record for dataset details and reuse information.

publicApr 2016View details →
zenodo20/100

NewsEye / READ AS training dataset from Austrian Newspapers (19th, early 20th C.)

<p>The dataset comprises Austrian newspaper pages from 19th and early 20th century with carefully annotated text. The page images were provided by the <a href="http://onb.ac.at/">Austrian National Library</a> and comprise 158 pages (training set). The data are formed according to the PAGE format (cf.&nbsp;Cf.&nbsp;<a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a>&nbsp;and the&nbsp;<a href="http://read.transkribus.eu/">READ </a>project. The guidelines with which the AS GT was created are uploaded here as well.</p>

restrictedApr 2021View details →
zenodo20/100

Newspaper Rock

Utah Source: Objaverse 1.0 / Sketchfab

opencc-byJun 2018View details →
zenodo12/100

The Portrayal of Residential Care in Newspapers: Considerations for Resident Involvement

<p>Interview transcripts of reporters, administrators, and long-term care residents to understand the considerations for residents to be involved in the news media production process.&nbsp;</p> <p>Data contain sensitive health information and are therefore not made publicly available.&nbsp;</p>

restrictedMay 2022View details →
zenodo12/100

Historic news comments from the G1 brazilian online newspaper

<p>This dataset contains data of user comments about the news published in the brazilian online newspaper G1&nbsp;from 2010 and before. It contains&nbsp;54,634 comments distributed between 4,549 news, which includes news published date,&nbsp;comments texts (in brazilian portuguese)&nbsp;and comments published date. That dataset intentionally ommits the titles, links, usernames or other kind of data that can redirect back to someones identity.</p>

restrictedOct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record