Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
82
datasets available to search
ShareScore release 0.9.0
Dataset results
82 results for “newspapers”
Dataset from "Merging Digital Humanities and Discourse Analysis in the Study of COVID-19 Vaccine Distribution in Norwegian Newspapers" (Sverdljuk et al. 2022)
<p>Contains URNs (identifiers) for the newspapers used in the corpus study "Merging Digital Humanities and Discourse Analysis in the Study of COVID-19 Vaccine Distribution in Norwegian Newspapers".</p> <p>For each subcorpus there is an Excel file containing references to the objects used, together with basic metadata.</p> <p>The corpus definitions can be used in various webapps of the DH-LAB at the National Library of Norway, e.g.:</p> <p><a href="https://beta.nb.no/dhlab/concordances/">https://beta.nb.no/dhlab/concordances/</a></p> <p><a href="https://beta.nb.no/dhlab/collocations/">https://beta.nb.no/dhlab/collocations/</a></p> <p>See more at <a href="https://www.nb.no/dh-lab/">https://www.nb.no/dh-lab/</a></p>
Viral Culture in Early Nineteenth-Century Europe newspaper dataset
<p>Dataset produced during the project Viral Culture in Early Nineteenth-Century Europe.</p> <p>The project traced text reuse by analysing large OCR'd newspaper collections using a BLAST based algorithm. This algorithm produces text clusters.</p> <p>This dataset contains two produced cluster datasets based on two different data collections.</p> <p>For the first dataset, the Austrian ANNO newspaper collection, this dataset contains metadata describing the used newspapers.</p> <p>For the second dataset, German-language newspapers in the Europeana collection, this dataset contains project produced metadata describing the newspapers used by the project, as well as the OCR's content for these newspaper issues. The OCR is produced with Tesseract OCR from digital page images downloaded from the Europeana services.</p> <p> </p>
Video of Using LOD to crowdsource Dutch WW2 underground newspapers on Wikipedia, SWIB2016, Bonn, 29-11-2016
<p><strong> </strong> Video registration of presentation about <a title="File:Using LOD to crowdsource Dutch WW2 underground newspapers on Wikipedia, SWIB2016, 29-11-2016.pdf" href="https://commons.wikimedia.org/wiki/File:Using_LOD_to_crowdsource_Dutch_WW2_underground_newspapers_on_Wikipedia,_SWIB2016,_29-11-2016.pdf">Using Linked Open Data to crowdsource Dutch WW2 underground newspapers on Wikipedia</a>, <a href="https://swib.org/swib16/programme.html" rel="nofollow">SWIB2016</a>, 29-11-2016 in Bonn.</p> <p>The video describes a project <a title="nl:Wikipedia:Wikiproject/Verzetskranten" href="https://nl.wikipedia.org/wiki/Wikipedia:Wikiproject/Verzetskranten">to systematically describe and interlink 1,300 Dutch underground newspapers from World War 2</a> on Wikipedia using linked open data.</p> <p>The project extracts contextual information about the newspapers from a book, converts it to structured data, and generates Wikipedia stubs linked to metadata, full texts, and each other.</p> <p>Volunteers are expanding the stubs into full articles, improving access to information about this historical period.</p> <div> <h2>Abstract</h2> </div> <p>During the second World War some 1.300 illegal newspapers were issued by the Dutch resistance. Right after the war as many of these newspapers as possible were physically preserved by Dutch memory institutions. They were described in formal library catalogues that were digitized and brought online in the 1990s. In 2010 the national collection of underground newspapers - some 200.000 pages - was full-text digitized in Delpher, the national aggregator for historical full-texts. Having created online metadata and full-texts for these publications, the third pillar <em>context</em> was still missing, making it hard for people to understand the historic background of the newspapers. We are currently running a project to tackle this contextual problem. We started by extracting contextual entries from a hard-copy standard work on Dutch illegal press and combined these with data from the library catalogue and Delpher into a central LOD triple store. We then created links between historically related newspapers and used Named Entity Recognition to find persons, organisations and places related to the newspapers. We further semantically enriched the data using DBPedia. Next, using an article template to ensure uniformity and consistency, we generated 1.300 Wikipedia article stubs from the database. Finally, we sought collaboration with the Dutch Wikipedia volunteer community to extend these stubs into full encyclopedic articles. In this way we can give every newspaper its own Wikipedia article, making these WW2 materials much more visible to the Dutch public, over 80% of whom uses Wikipedia. At the same time the triple store can serve as a source for alternative applications, like data visualizations. This will enable us to visualize connections and networks between underground newspapers, as they developed over time between 1940 and 1945.</p> <div> <h2>Presentation slides</h2> </div> <ul> <li><a title="File:Using LOD to crowdsource Dutch WW2 underground newspapers on Wikipedia, SWIB2016, 29-11-2016.pdf" href="https://commons.wikimedia.org/wiki/File:Using_LOD_to_crowdsource_Dutch_WW2_underground_newspapers_on_Wikipedia,_SWIB2016,_29-11-2016.pdf">Using LOD to crowdsource Dutch WW2 underground newspapers on Wikipedia</a></li> <li><a href="https://swib.org/swib16/slides/janssen_using_lod.pdf" rel="nofollow">https://swib.org/swib16/slides/janssen_using_lod.pdf</a></li> <li><a href="https://doi.org/10.5281/zenodo.13132987" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13132987</a></li> </ul>
Newspaper references to Irish language Court Interpreters 1796-1922
<p>This file consists of a corpus made up of transcriptions of newspaper extracts that mention Irish language court interpreters. Much of the information was used in <em>Irish Speakers, Interpreter and the Courts</em> (2019) by Mary Phelan, published by Four Courts Press and the Irish Legal History Society.</p>
NewsEye / READ OCR training dataset from Austrian Newspapers (19th C.)
<p>The dataset comprises Austrian newspaper pages from 19th and early 20th century with carefully corrected text. The page images were provided by the <a href="http://onb.ac.at/">Austrian National Library</a> and comprise 148 pages (training set) and 13 pages (validation set). The data are formed according to the PAGE format (cf. Cf. <a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a> and the <a href="http://read.transkribus.eu/">READ </a>project.</p>
Bibliographic metadata for Khalīl al-Khūrī's weekly newspaper Ḥadīqat al-Akhbār (حديقة الاخبار) published by in Beirut (1858–65)
<p>This repository contains bibliographic metadata for the first 357 issues of the weekly newspaper <em>Ḥadīqat al-Akhbār</em> published by Khalīl al-Khūrī in Beirut between 1858 and 1865. Digital facsimiles are not publicly available online. Links to ocal facsimiles have been added from #142 onwards.</p> <p>The repository contains machine-actionable bibliographic metadata on the issue level as BibTeX, MODS XML, and Zotero RDF. I produced bootstrapped TEI XML editions for each issue, which link to local copies of facsimiles (not part of this data set).</p> <p>This data set is part of Open Arabic Periodical Editions (<a href="https://openarabicpe.github.io">OpenArabicPE</a>).</p>
A transnational newspaper dataset covering Spenceanism
<p>To analyse the media legacy of English radical thinker Thomas Spence (1750–1814), his revolutionary "Plan", and his disciples (the "Spencean Philanthropists") in global press circuits, a dataset consisting of 275 Spencean-related articles in newspapers from Ireland (106 articles), the British West Indies (35), British India (29), the Australian colonies (4), Canada (1), and the United States (100) was created. The corpus consists of either the full text of each article or relevant extracts, as well as additional metadata such as source, date, title (if applicable), keyword, and region. In particular, the following databases have been relied upon: the Irish Newspaper Archives for Ireland, the Caribbean Newspapers: Digital Library of the Caribbean and Caribbean Newspapers 1718–1876 for the British West Indies, Newspapers & Gazettes – Trove for Australia, America's Historical Newspapers and Chronicling America: Historic American Newspapers for the US, Newspapers.com by Ancestry for Canada, Ireland, and the US, and the British Newspaper Archive for British India, Ireland, and the Caribbean. These databases were searched using the following keywords: "Thomas Spence" [1750–1814], "Spence's Plan", "Spencean" / "Spenceans", and "Spenceanism"; "swinish multitude", "people's farm", and "pigs' meat" were also used. All databases checked for Australia, the Caribbean, India, and Canada have been exhausted, and many Irish and US-American articles have also been grabbed. Due to the low scan quality of databases, most articles needed to be copied and corrected manually. The 275 articles of the corpus consist of 167,515 tokens. In addition, 157 articles on Spence and the Spenceans from British newspapers have been downloaded, too, and were added to this dataset for qualitative analysis and comparative purposes. Overall, the dataset made available here thus consists of 432 articles. The results from analysing this corpus are shown and discussed in a paper entitled "Transnational Echoes of Spenceanism: A Text Mining Exploration in English-Language Newspapers (1790–1850)", which is accepted for publication in the International Review of Social History (IRSH).</p>
Dataset and Models for Detection of News Agency Releases in Historical Newspapers
<p>This record contains the annotated datasets and models used and produced for the work reported in the Master Thesis "<em>Where Did the News come from? Detection of News Agency Releases in Historical Newspapers</em> " (<a href="https://infoscience.epfl.ch/record/305129?&ln=en">link</a>).</p> <p>Please cite this report if you are using the models/datasets or find it relevant to your research:</p> <pre><code>@article{Marxen:305129, title = {Where Did the News Come From? Detection of News Agency Releases in Historical Newspapers}, author = {Marxen, Lea}, pages = {114p}, year = {2023}, url = {http://infoscience.epfl.ch/record/305129}, }</code></pre> <p><br> <strong>1. DATA</strong></p> <p>The <strong>newsagency-dataset</strong> contains historical newspaper articles with annotations of news agency mentions. The articles are divided into French (fr) and German (de) subsets and a train, dev and test set respectively. The data is annotated at token-level in the CoNLL format with IOB tagging format.</p> <p>The distribution of articles in the different sets is as follows:</p> <table> <caption>Dataset Statistics</caption> <thead> <tr> <th scope="row"> </th> <th scope="col">Lg.</th> <th scope="col">Docs</th> <th scope="col">Agency Mentions</th> </tr> </thead> <tbody> <tr> <th scope="row">Train</th> <td>de</td> <td>333</td> <td>493</td> </tr> <tr> <th scope="row"> </th> <td>fr</td> <td>903</td> <td>1,122</td> </tr> <tr> <th scope="row">Dev</th> <td>de</td> <td>32</td> <td>26</td> </tr> <tr> <th scope="row"> </th> <td>fr</td> <td>110</td> <td>114</td> </tr> <tr> <th scope="row">Test</th> <td>de</td> <td>32</td> <td>58</td> </tr> <tr> <th scope="row"> </th> <td>fr</td> <td>120</td> <td>163</td> </tr> </tbody> </table> <p>Due to an error, there are seven duplicated articles in the French test set (<em>article IDs: courriergdl-1847-10-02-a-i0002, courriergdl-1852-02-14-a-i0002, courriergdl-1860-10-31-a-i0016, courriergdl-1864-12-15-a-i0005, lunion-1860-11-27-a-i0004, lunion-1865-02-05-a-i0012, lunion-1866-02-16-a-i0009</em>).</p> <p> </p> <p><strong>2. MODELS</strong></p> <p>The two agency detection and classification models used for the inference on the <em><a href="https://impresso-project.ch/">impresso</a> </em>Corpus are released as well:</p> <ul> <li><strong>newsagency-model-de</strong>: based on <a href="https://www.deepset.ai/german-bert">German BERT</a> (with maximum sequence length 128), fine-tuned with the German training set of the newsagency-dataset</li> <li><strong>newsagency-model-fr</strong>: based on <a href="https://huggingface.co/dbmdz/bert-base-french-europeana-cased">French Europeana BERT</a> (with maximum sequence length 128), fine-tuned with the French training set of the newsagency-dataset</li> </ul> <p>The models perform multitask classification with two prediction heads, one for token-level agency entity classification and one for sentence-level (<em>has_agency: yes/no</em>). They can be run with TorchServe, for details see the <a href="https://github.com/impresso/newsagency-classification/tree/main/lib/bert_classification">newsagency-classification</a> repository.</p> <p> </p> <p>Please refer to the report for further information or contact us.</p> <p> </p> <p><strong>3. CODE</strong></p> <p><a href="https://github.com/impresso/newsagency-classification">https://github.com/impresso/newsagency-classification</a></p> <p> </p> <p><strong>4. CONTACT</strong></p> <p>Maud Ehrmann (EPFL-DHLAB)<br> Emanuela Boros (EPFL-DHLAB)</p>
Newspaper review and analysis data on media coverage of budget issues in Nigeria
<p>The data was collected from analysis of six newspaper publications to establish citizen engagement and media coverage of the budget discourse from 2009 to 2013 as part of the investigation of the use of the online national budget of Nigeria. The work was part of the 'Exploring the Emerging Impacts of Open Data in Developing Countries' (ODDC) research project.</p>
Newspaper Box
This Scan is part of the Urban Relics Project. Documenting the disapearing street landmarks of New York City. Source: Objaverse 1.0 / Sketchfab
NewsEye / READ AS training dataset from Finnish Newspapers (19th C.)
<p>The dataset comprises finnish newspaper pages from 19th century with carefully annotated text. The page images were provided by the <a href="https://www.kansalliskirjasto.fi/en/">National Library Finland</a> (NLF) and comprise 200 pages (training set). The data are formed according to the PAGE format (cf. Cf. <a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a> and the <a href="http://read.transkribus.eu/">READ </a>project. The guidelines with which the AS GT was created are uploaded here as well.</p>
PT2vec - A Portuguese text corpus created from online newspapers
<p>A Portuguese text corpus created from online newspapers with a total of 394,825,480 tokens and 33,089,734 sentences.</p> <p>If you use this repository, please cite this paper:</p> <p>Pinto JP, Viana P, Teixeira I, Andrade M. 2022. Improving word embeddings in Portuguese: increasing accuracy while reducing the size of the corpus. PeerJ Computer Science 8:e964 <a href="https://doi.org/10.7717/peerj-cs.964">https://doi.org/10.7717/peerj-cs.964</a></p>
Cheltenham Town 1933-1934 newspaper archive
<p>Newspaper articles reporting on Cheltenham Town A.F.C. for the football season of 1933-1934. Articles cover the dates of 1933-06-01 to 1934-05-31. Articles are available as HTML files as well as a mix of jpeg and png scans of original articles and pictures.</p> <p>RIS file is available for import into reference management software.</p> <p>XML file available containing text of articles.</p> <p>The contents are considered to be public domain as it is over 70 years after the end of the year in which the work was created or first published.</p>
Cheltenham Town 1932-1933 newspaper archive
<p>Newspaper articles reporting on Cheltenham Town A.F.C. for the football season of 1932-1933. Articles cover the dates of 1932-04-01 to 1933-05-31. Articles are available as HTML files as well as a mix of jpeg and png scans of original articles and pictures.</p> <p>RIS file is available for import into reference management software.</p> <p>XML file available containing text of articles.</p> <p>The contents are considered to be public domain as it is over 70 years after the end of the year in which the work was created or first published.</p>
Газеты | The Object soviet newspaper props
Советские газеты для моего проекта. Делал для себя, но можно пользоваться всем, кто хочет. Тут только текстуры, поэтому модели и формы делайте уже сами. Если вы вдруг хотите мне как-то помочь в развитии канала, то я не против ( НО НЕ НАСТАИВАЮ!): https://www.donationalerts.com/r/gradislav Не забывайте на мой ютуб заходить, жду от вас подписочки и комментов! https://www.youtube.com/c/NOQUALITY96/videos Source: Objaverse 1.0 / Sketchfab
Historical Newspapers Ground Truth
<p>This dataset contains 50 pages of ground truth data for digitized historical newspapers from the Berlin State Library for training and evaluation of OCR/OLR systems as produced in the context of the EU ICT-PSP project Europeana Newspapers (http://www.europeana-newspapers.eu/).</p> <p>The dataset comprises of the following resources:</p> <ul> <li><strong>gt_page.zip </strong>Ground Truth files in PAGE-XML format (cf. https://github.com/PRImA-Research-Lab/PAGE-XML)</li> <li><strong>img_full.zip </strong>Full resolution scanned images in TIF format</li> <li><strong>img_bin.zip</strong> Binarized (using the Gatos method) images in TIF format</li> <li><strong>ocr_full.zip</strong> OCR (FineReaderEngine11) results for full resolution images</li> <li><strong>ocr_bin.zip</strong> OCR (FineReaderEngine11) results for binarized images</li> </ul>
Word2Vec Models Dutch Newspapers
<p>Word Embedding models trained on 6 national Dutch newspapers. </p> <p>We use the Gensim implementation of Word2Vec to train four embedding models per newspaper, each representing one decade between 1950 and 1990. The models were trained using C-BOW with hierarchical softmax, with a dimensionality of 300, a minimal word count and context of 5, and downsampling of 10<sup>-5</sup></p> <p>These models belong to the article: Using Word Embeddings to Examine Gender Bias in Dutch Newspapers, 1950-1990</p>
Estonian Historical Newspaper Crowdsourced OCR Corrections
<div> <div>This dataset consists of newspaper articles from the National Library of Estonia's DIGAR archive and their respective crowdsourced corrections.</div> </div>
Newspaper valence datasets for van der Veen & Bleich (2024)
<p>Replication datasets for average valence calculations (full newspapers and national representative corpora).</p>
Clusters of topic modelling and naive Bayes classifier - FR CH newspapers
<p>Clusters of articles based on annotations produced by topic modelling and naive Bayes classifier applied to French language newspapers of Switzerland, published between 1900 and 1944 and containing the characters "europ", extracted from the impresso app. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.