Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

82

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

82 results for “Newspapers”

Learn how ShareScore rates datasets ↗
zenodo48/100

Models for "A data-driven approach to studying changing vocabularies in historical newspaper collections"

<p>NOTE: This is a badly rendered version of the README within the archive.</p> <p><strong>A data-driven approach to studying changing vocabularies in historical newspaper collections</strong></p> <p>Simon Hengchen,* Ruben Ros,** Jani Marjanen,*** Mikko Tolonen***</p> <p>*<a href="https://spraakbanken.gu.se/en/about/staff/simon">Spr&aring;kbanken Text</a>, University of Gothenburg, Sweden and <a href="https://iguanodon.ai">iguanodon.ai</a>, Belgium: firstname.lastname@gu.se<br> **<a href="https://www.c2dh.uni.lu/people/ruben-ros">Centre for Contemporary and Digital History (C2DH)</a>, University of Luxembourg:&nbsp;firstname.lastname@uni.lu<br> ***<a href="https://www.helsinki.fi/en/researchgroups/computational-history">COMHIS</a>, University of Helsinki:&nbsp;<a href="mailto:firstname.lastname@helsinki.fi">firstname.lastname@helsinki.fi</a>;</p> <p>These are the supplementary materials for the DH2019 paper&nbsp;<em>A data-driven approach to the changing vocabulary of the &lsquo;nation&rsquo; in English, Dutch, Swedish and Finnish newspapers, 1750-1950</em>, as well as the 2021 Digital Scholarship in the Humanities publication available in OpenAccess: <a href="https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793">https://academic.oup.com/dsh/article/36/Supplement_2/ii109/6421793</a>. If you end up using whole or parts of this resource, please use the following citation(s):</p> <ul> <li>Hengchen, S., Ros, R., and Marjanen, J. (2019). A data-driven approach to the changing vocabulary of the &#39;nation&#39; in English, Dutch, Swedish and Finnish newspapers, 1750-1950. In&nbsp;<em>Proceedings of the Digital Humanities (DH) conference 2019, Utrecht, The Netherlands</em></li> </ul> <p>and/or:</p> <ul> <li>Hengchen, S., Ros, R., Marjanen, J. and Tolonen, M., 2021. A data-driven approach to studying changing vocabularies in historical newspaper collections. Digital Scholarship in the Humanities, 36(Supplement_2), pp.ii109-ii126.</li> </ul> <p>or alternatively use one of the following&nbsp;<code>bib</code>s:</p> <pre><code>@inproceedings{hengchen2019nation, title="A data-driven approach to the changing vocabulary of the 'nation' in {E}nglish, {D}utch, {S}wedish and {F}innish newspapers, 1750-1950.", author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani}, year={2019}, address = "Utrecht, The Netherlands", booktitle={Proceedings of the Digital Humanities (DH) conference 2019} }</code></pre> <pre><code>@article{hengchen2021data, title={A data-driven approach to studying changing vocabularies in historical newspaper collections}, author={Hengchen, Simon and Ros, Ruben and Marjanen, Jani and Tolonen, Mikko}, journal={Digital Scholarship in the Humanities}, volume={36}, number={Supplement\_2}, pages={ii109--ii126}, year={2021}, publisher={Oxford University Press} }</code></pre> <p>&nbsp;</p> <p>Files</p> <p>This archive contains two folders -- one per diachronic representation method -- as well as this README. The folders each contain four folders, which contain the models for their respective languages. As can be inferred from the small datasize, most of the earlier models are not reliable and should not be used, but are still made available. This work is licensed under a&nbsp;<a href="http://creativecommons.org/licenses/by-sa/4.0/">Creative Commons Attribution-ShareAlike 4.0 International License</a>.</p> <p><strong>Source material</strong></p> <p>Finnish:</p> <p>The models were created with data from the Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland (National Library of Finland, 2011). We used everything in the corpus.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h fi* 12M fi_1820_SGNS_corpus_file.gensim 89M fi_1840_SGNS_corpus_file.gensim 797M fi_1860_SGNS_corpus_file.gensim 7.0G fi_1880_SGNS_corpus_file.gensim 22G fi_1900_SGNS_corpus_file.gensim</code></pre> <p>Swedish:</p> <p>The models were created with data from the Kubhist 2 corpus (Spr&aring;kbanken) -- more precisely, the data dumps available at&nbsp;<a href="https://spraakbanken.gu.se/lb/resurser/meningsmangder/">https://spraakbanken.gu.se</a>. After a manual evaluation of Swedish embeddings trained without pre-processing seemed to show that the embeddings were of low quality, we retrained models, only keeping sentences that were at least 10 tokens long and were constituted of at least 50% of lemmas as per the KORP processing pipeline (Borin et al, 2012).</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h sv* 1.6M sv_1740_SGNS_corpus_file.gensim 44M sv_1760_SGNS_corpus_file.gensim 124M sv_1780_SGNS_corpus_file.gensim 228M sv_1800_SGNS_corpus_file.gensim 678M sv_1820_SGNS_corpus_file.gensim 1.6G sv_1840_SGNS_corpus_file.gensim 4.5G sv_1860_SGNS_corpus_file.gensim 6.5G sv_1880_SGNS_corpus_file.gensim 113M sv_1900_SGNS_corpus_file.gensim</code></pre> <p>Dutch:</p> <p>The models were created with data from the Delpher newspaper archive (Royal Dutch Library, 2017), through data dumps for newspapers until and including 1876, and through API hits for articles from 1877 to 1899 (included).</p> <ul> <li>For anything pre-1877 we discarded full texts that had, in the metadata, anything else than exclusively&nbsp;<code>nl</code>&nbsp;or&nbsp;<code>NL</code>&nbsp;as a language tag.</li> <li>For the full texts between 1877 and 1899: we queried the API for all items in the &ldquo;artikel&rdquo; category that contained the determiner&nbsp;<code>de</code>.</li> </ul> <p>Our assumption was that most articles should contain&nbsp;<code>de</code>&nbsp;at least once, and those that didn&#39;t were too short to be deemed interesting. A subsequent study showed that was not exactly the case, but we were reassured by the fact that left-out articles were probably &quot;shipping or financial reports&quot; (thanks go to Melvin Wevers). We also did not include the colonial newspapers for our embeddings. This is motivated by our research questions. A list of removed newspapers is available on request.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h nl* 6.8M nl_1620_SGNS_corpus_file.gensim 7.9M nl_1640_SGNS_corpus_file.gensim 43M nl_1660_SGNS_corpus_file.gensim 78M nl_1680_SGNS_corpus_file.gensim 138M nl_1700_SGNS_corpus_file.gensim 243M nl_1720_SGNS_corpus_file.gensim 287M nl_1740_SGNS_corpus_file.gensim 431M nl_1760_SGNS_corpus_file.gensim 825M nl_1780_SGNS_corpus_file.gensim 1.2G nl_1800_SGNS_corpus_file.gensim 1.8G nl_1820_SGNS_corpus_file.gensim 3.1G nl_1840_SGNS_corpus_file.gensim 5.2G nl_1860_SGNS_corpus_file.gensim 13G nl_1880_SGNS_corpus_file.gensim</code></pre> <p>English:</p> <p>The models were created with data from the British Library Newspapers collection (<a href="https://www.gale.com/intl/primary-sources/british-library-newspapers%5D">link</a>), the Nichols collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-burney-newspapers-collection">link</a>), and the Burney collection (<a href="https://www.gale.com/intl/c/17th-and-18th-century-nichols-newspapers-collection">link</a>). We used everything in the corpora. For English, only SGNS_ALIGN models are available. We thank Gale Cengage for their help with this project.</p> <p>Filesizes:</p> <pre><code>[simon@taito-login3 SGNS]$ du -h en* 4.3M en_1620_SGNS_corpus_file.gensim 11M en_1640_SGNS_corpus_file.gensim 11M en_1660_SGNS_corpus_file.gensim 106M en_1680_SGNS_corpus_file.gensim 409M en_1700_SGNS_corpus_file.gensim 1.7G en_1720_SGNS_corpus_file.gensim 834M en_1740_SGNS_corpus_file.gensim 2.4G en_1760_SGNS_corpus_file.gensim 5.3G en_1780_SGNS_corpus_file.gensim 5.5G en_1800_SGNS_corpus_file.gensim 15G en_1820_SGNS_corpus_file.gensim 42G en_1840_SGNS_corpus_file.gensim 65G en_1860_SGNS_corpus_file.gensim 88G en_1880_SGNS_corpus_file.gensim 26G en_1900_SGNS_corpus_file.gensim 21G en_1920_SGNS_corpus_file.gensim 6.3G en_1940_SGNS_corpus_file.gensim</code></pre> <p><strong>Word embeddings</strong></p> <p>For every language, we train diachronic embeddings as follows. We divide the data in 20-year time bins. We train SGNS_UPDATE and SGNS_ALIGN models. Current research on German (Schlechtweg et al, 2019) and English (Shoemark et al, 2019) indicates you should use the SGNS_ALIGN models.&nbsp;<strong>For EN, FI, NL, no tokens (including punctuation) were removed nor altered, aside from lowercasing</strong>. For SV, see above. Parameters are as follows: SGNS architecture (Mikolov et al 2013), window size of 5, frequency threshold of 100, 5 epochs, 300 dimensions (or 100 for EN).</p> <ul> <li>For SGNS_UPDATE: We first train a model for the first time bin&nbsp;<code>t</code>. To train the model for&nbsp;<code>t+1</code>, we use the&nbsp;<code>t</code>&nbsp;model to initialise the vectors for&nbsp;<code>t+1</code>, set the learning rate to correspond to the end learning rate of&nbsp;<code>t</code>, and continue training. This approach, closely following Kim et al (2014), has the advantage of avoiding the need for post-training vector space alignment.</li> </ul> <p>The Python snippet below, which makes use of gensim (Rehurek and Sojka, 2010), illustrates the approach. Special thanks go to Sara Budts.</p> <pre><code>## dict_files[key] is a dictionary with double decades as keys and a corresponding LineSentence object as value: https://radimrehurek.com/gensim/models/word2vec.html#gensim.models.word2vec.LineSentence count = 0 for key in sorted(list(dict_files.keys())): if count == 0: ## This is the first model. model = gensim.models.Word2Vec(corpus_file=dict_files[key], min_count=100, sg=1 ,size=300, workers=64, seed=1830, iter=5) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) print("Model saved, on to the next\n") count += 1 if count &gt; 0: ## this is for the subsequent models. print("model for double decade starting in",str(key)) model = gensim.models.Word2Vec.load(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin-20)+".w2v")) print("previous model loaded") model.build_vocab(corpus_file=dict_files[key], update=True) model.train(corpus_file=dict_files[key], total_words = model.corpus_count, total_examples = model.corpus_count, start_alpha = model.alpha, end_alpha = model.min_alpha, epochs=model.epochs) model.save(os.path.join(data_path_final,"KIM",lang+"_"+str(timebin)+".w2v")) </code></pre> <ul> <li>For SGNS_ALIGN: We independently train models for all time bins. The models in this repository are&nbsp;<em>NOT</em>&nbsp;aligned, leaving you the choice of how to align them. For example,&nbsp;<a href="https://gist.github.com/quadrismegistus/09a93e219a6ffc4f216fb85235535faf">here</a>&nbsp;is a link to code by Ryan Heuser to do just that. Models were trained with the&nbsp;<code>count == 0</code>&nbsp;scenario in the snippet above.</li> </ul> <p><strong>Acknowledgments</strong></p> <p>This work has been supported by the European Union&#39;s Horizon 2020 research and innovation programme under grant 770299&nbsp;<a href="https://www.newseye.eu/">NewsEye</a>. Specials thanks go to the data providers/collection-holding institutions: the Finnish Language Bank, the Swedish Language Bank, the Royal Dutch Library, and Gale Cengage.</p> <p>The authors would like to thank the following persons and group, listed alphabetically: Antoine Doucet, Antti Kanner, Axel-Jean Caurant, Dominik Schlechtweg, Eetu M&auml;kel&auml;, Elaine Zosa, Estelle Bunout, Haim Dubossarsky, Joris van Eijnatten, Krister Lind&eacute;n, Lars Borin, Lidia Pivovarova, Melvin Wevers, Nina Tahmasebi, Sara Budts, Senka Drobac, Tanja S&auml;ily, the COMHIS group, and Steven Claeyssens. Computational resources were provided by CSC &ndash; IT Center for Science Ltd.</p> <p><strong>References</strong></p> <p>Borin, L., Forsberg, M., Roxendal, J. (2012). Korp-the corpus infrastructure of Spr&auml;kbanken,in: LREC. pp. 474&ndash;478.</p> <p>Kim, Y., Chiu, Y.I., Hanaki, K., Hegde, D. and Petrov, S. (2014). Temporal Analysis of Language through Neural Language Models.&nbsp;<em>ACL 2014</em>, p.61.</p> <p>Mikolov, T., Chen, K., Corrado, G. and Dean, J. (2013). Efficient estimation of word representations in vector space.&nbsp;<em>arXiv preprint arXiv:1301.3781</em>.</p> <p>National Library of Finland (2011).&nbsp;<em>The Finnish Sub-corpus of the Newspaper and Periodical Corpus of the National Library of Finland, Kielipankki Version</em>&nbsp;[text corpus]. Kielipankki. Retrieved from&nbsp;<a href="http://urn.fi/urn:nbn:fi:lb-2016050302">http://urn.fi/urn:nbn:fi:lb-2016050302</a>.</p> <p>Rehurek, R. and Sojka, P. (2010). Software framework for topic modelling with large corpora. In&nbsp;<em>Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks</em>.</p> <p>Royal Dutch Library (2017).&nbsp;<em>Delpher open krantenarchief (1.0)</em>. Den Haag, 2017.</p> <p>Schlechtweg D., H&auml;tty A, del Tredici M., and Schulte im Walde S. (2019). A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains. In&nbsp;<em>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, Florence, Italy. ACL.</p> <p>Shoemark, P., Liza, F.F., Nguyen, D., Hale, S. and McGillivray, B. (2019). Room to Glo: A Systematic Comparison of Semantic Change Detection Approaches with Word Embeddings. In&nbsp;<em>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 66-76)</em>, Hong Kong.</p> <p>Spr&aring;kbanken.&nbsp;<em>The Kubhist Corpus</em>. Department of Swedish, University of Gothenburg.&nbsp;<a href="https://spraakbanken.gu.se/korp/?mode=kubhist">https://spraakbanken.gu.se/korp/?mode=kubhist</a>.</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Diachronic word embeddings from 19th-century newspapers digitised by the British Library (1800-1919)

<p>Word vectors related to the paper&nbsp;<em>Machines in the media: semantic change in the lexicon&nbsp;of mechanization in 19th-century British newspapers&nbsp;</em>by Nilo Pedrazzini and Barbara McGillivray (2022).</p> <p>The embeddings were trained on a 4.2-billion-word corpus of 19th-century British newspapers using Word2Vec and the following parameters:</p> <pre><code>sg = True min_count = 1 window = 3 vector_size = 200 epochs = 5</code></pre> <p>The embeddings&nbsp;are divided into periods of ten years each, with the vectors from each decade aligned to the ones from the most recent decade (1910s) using Orthogonal Procrustes.</p> <p>See related GitHub repository for the full documentation:&nbsp;<a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a></p> <p>Project webpage (Living with Machines):&nbsp;<a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p>

opencc-by-4.0Oct 2022View details →
zenodo48/100

Survey of digitized newspaper interfaces (dataset and notebooks)

<p>This record contains the datasets and jupyter notebooks which support the analysis presented in the paper &quot;Historical Newspaper User Interfaces: A Review&quot;. Please refer to the paper or the github repository for more information (see links below), or do not hesitate to contact us!</p>

opencc-by-4.0Aug 2019View details →
zenodo48/100

trove-newspaper-issues

<p>This dataset contains information about the published issues of newspapers digitised and made available through Trove. The data was harvested from the Trove API, using <a href="https://glam-workbench.net/trove-newspapers/harvest_newspaper_issues/">this notebook in the GLAM Workbench</a>.</p> <p>There are two data files:</p> <ul> <li><code>newspaper_issues_totals_by_year.csv</code> &ndash; the total number of newspaper issues per year for each digitised newspaper</li> <li><code>newspaper_issues.csv</code> &ndash; a complete list of newspaper issues available from Trove</li> </ul> <h2>newspaper_issues_totals_by_year.csv</h2> <p>The dataset contains the following columns:</p> <table> <tbody> <tr> <td><strong>Column</strong></td> <td><strong>Contents</strong></td> </tr> <tr> <td><code>title</code></td> <td>newspaper title</td> </tr> <tr> <td><code>title_id</code></td> <td>newspaper id</td> </tr> <tr> <td><code>state</code></td> <td>place of publication</td> </tr> <tr> <td><code>year</code></td> <td>year published</td> </tr> <tr> <td><code>issues</code></td> <td>number of issues</td> </tr> </tbody> </table> <h2>newspaper_issues.csv</h2> <p>The dataset contains the following columns:</p> <table> <tbody> <tr> <td><strong>Column</strong></td> <td><strong>Contents</strong></td> </tr> <tr> <td><code>title</code></td> <td>newspaper title</td> </tr> <tr> <td><code>title_id</code></td> <td>newspaper id</td> </tr> <tr> <td><code>state</code></td> <td>place of publication</td> </tr> <tr> <td><code>issue_id</code></td> <td>issue identifier</td> </tr> <tr> <td><code>issue_date</code></td> <td>date of publication (YYYY-MM-DD)</td> </tr> </tbody> </table> <p>To keep the file size down, I haven't included an <code>issue_url</code> in this dataset, but these are easily generated from the <code>issue_id</code>. Just add the <code>issue_id</code> to the end of <code>http://nla.gov.au/nla.news-issue</code>. For example: <a href="http://nla.gov.au/nla.news-issue495426">http://nla.gov.au/nla.news-issue495426</a>. Note that when you follow an issue url, you actually get redirected to the url of the first page in the issue.</p>

opencc-zeroOct 2021View details →
zenodo48/100

Decade-level Word2Vec models from automatically transcribed 19th-century newspapers digitised by the British Library (1800-1919)

<p>Word embeddings trained on a 4.2-billion-word corpus of 19th-century British newspapers using Word2Vec and the following parameters:</p> <pre><code>sg = True min_count = 5 window = 5 vector_size = 100 epochs = 5</code></pre> <p>The embeddings&nbsp;are divided into periods of ten years each. Unlike those in <a href="https://doi.org/10.5281/zenodo.7181682">this repository</a>, these were not aligned and OCR errors skimmed from the vocabulary.&nbsp;</p> <p>See related GitHub repository for the full documentation:&nbsp;<a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a></p> <p>Project website (Living with Machines):&nbsp;<a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p>

opencc-by-4.0May 2023View details →
zenodo44/100

Datasets and Models for Historical Newspaper Article Segmentation

<p>This record contains the datasets and models used and produced for the work reported in the paper &quot;<em>Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers</em>&quot; (<a href="https://infoscience.epfl.ch/record/282863?ln=en">link</a>).</p> <p>Please cite this paper if you are using the models/datasets or find it relevant to your research:</p> <pre><code>@article{barman_combining_2020, title = {{Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers}}, author = {Raphaël Barman and Maud Ehrmann and Simon Clematide and Sofia Ares Oliveira and Frédéric Kaplan}, journal= {Journal of Data Mining \&amp; Digital Humanities}, volume= {HistoInformatics} DOI = {10.5281/zenodo.4065271}, year = {2021}, url = {https://jdmdh.episciences.org/7097}, }</code></pre> <p><br> <strong>Please note that this record contains data under different licenses.</strong><br> <br> <strong>1. DATA</strong></p> <ul> <li><strong>Annotations (json files)</strong>: JSON files contains image annotations, with one file per newspaper containing region annotations (label and coordinates) in VIA format. The following licenses apply: <ul> <li>&nbsp;luxwort.json: those annotations are under a <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 license</a>. Please refer to the right statement specified for each image in the file.</li> <li>GDL.json, IMP.json and JDG.json: those annotations are under a <a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">CC BY-SA 4.0 license</a>.</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li><strong>Image files: </strong>The archive images.zip contains the Swiss titles image files (GDL, IMP, JDG) used for the experiments described in the paper. Those images are under copyright (property of the journal <em>Le Temps </em>and of <em>ArcInfo</em>) and can be used <em>for academic research or educational purposes only</em>. Redistribution, publication or commercial use are not permitted. These terms of use are similar to the following right statement: <a href="http://rightsstatements.org/vocab/InC-EDU/1.0/">http://rightsstatements.org/vocab/InC-EDU/1.0/</a></li> </ul> <p>&nbsp;</p> <p><strong>2. MODELS</strong></p> <p>Some of the best models are released under a <a href="https://creativecommons.org/licenses/by-sa/4.0/">CC BY-SA 4.0</a> license (they are also available as assets of the current Github <a href="https://github.com/dhlab-epfl/dhSegment-text/releases/tag/0.1">release</a>).</p> <ul> <li><strong>JDG_flair-FT</strong>: this model was trained on JDG using french Flair and FastText embeddings. It is able to predict the four classes presented in the paper (<code>Serial</code>, <code>Weather</code>, <code>Death notice</code> and <code>Stocks</code>).</li> <li><strong>Luxwort_obituary_flair-bpemb</strong>: this model was trained on Luxwort using multilingual Flair and Byte-pair embeddings. It is able to predict the <code>Death notice</code> class.</li> <li><strong>Luxwort_obituary_flair-FT_indomain</strong>: this model was trained on Luxwort using in-domain Flair and FastText embeddings (trained on Luxwort data). It is also able to predict the <code>Death notice</code> class.</li> </ul> <p>Those models can be used to predict probabilities on new images using the same code as in the original <a href="https://github.com/dhlab-epfl/dhSegment">dhSegment</a> repository. One needs to adjust three parameters to the <code>predict</code> function: 1) <code>embeddings_path</code> (the path to the embeddings list), 2) <code>embeddings_map_path</code>(the path to the compressed embedding map), and 3) <code>embeddings_dim</code> (the size of the embeddings).</p> <p>Please refer to the paper for further information or contact us.</p> <p>&nbsp;</p> <p><strong>3. CODE:&nbsp;</strong></p> <p><a href="https://github.com/dhlab-epfl/dhSegment-text">https://github.com/dhlab-epfl/dhSegment-text</a></p> <p><br> <strong>4. ACKNOWLEDGEMENTS</strong><br> We warmly thank the journal <a href="https://letemps.ch">Le Temps</a> (owner of <em>La Gazette de Lausanne</em> and the <em>Journal de Gen&egrave;ve</em>) and the group <a href="https://www.arcinfo.ch/">ArcInfo</a> (owner of <em>L&#39;Impartial</em>) for accepting to share the related datasets for academic purposes. We also thank the <a href="https://bnl.public.lu/fr.html">National Library of Luxembourg</a> for its support with all steps related to the <em>Luxemburger Wort</em> annotation release.<br> This work was realized in the context of the <a href="https://impresso-project.ch"><em>impresso</em> - Media Monitoring of the Past</a> project and supported by the Swiss National Science Foundation under grant CR- SII5_173719.<br> <br> <strong>5. CONTACT</strong><br> Maud Ehrmann (EPFL-DHLAB)<br> Simon Clematide (UZH)</p>

openother-ncJan 2021View details →
zenodo44/100

Gado2: multilingual newspapers from the Netherlands Indies

<p>This Handwritten Text Recognition (HTR) xml-page file dataset contains the ground truths of the Gado2 named entity processing application for newspapers from the Netherlands Indies and Indonesia, see: https://github.com/KBNLresearch/gado2. Optical Character Recognition (OCR) resulted in high Character Error Rates (CER) due to the inferior quality of many scans. In contrast, HTR led to CERs below 0.5 percent thus increasing the efficiency of the NER engine. All uploaded files are free of errors and fully tagged. A relevant knowledge base of Indonesian persons, places and organisations is attached in json format for entity linking.</p>

opencc-by-4.0May 2021View details →
zenodo44/100

NewsEye / READ AS training dataset from French Newspapers (19th, early 20th C.)

<p>The dataset comprises French newspaper pages from 19th and early 20th century with annotated text. The page images were provided by the&nbsp;<a href="https://www.bnf.fr/en">French National Library</a> and comprise 183 pages (training set). The data are formed according to the PAGE format (cf.&nbsp;Cf.&nbsp;<a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a>&nbsp;and the&nbsp;<a href="http://read.transkribus.eu/">READ </a>project. The guidelines with which the AS GT was created are uploaded here as well.</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

19th Century United States Newspaper images predicted as Photographs with labels for "human", "animal", "human-structure" and "landscape"

<p>The Dataset contains images derived from the Newspaper Navigator (news-navigator.labs.loc.gov/), a dataset of images drawn from the Library of Congress Chronicling America collection (<a href="https://chroniclingamerica.loc.gov/">chroniclingamerica.loc.gov/</a>).&nbsp;</p> <blockquote> <p>[The Newspaper Navigator dataset] consists of extracted visual content for 16,358,041 historic newspaper pages in&nbsp;<em>Chronicling America</em>. The visual content was identified using an object detection model trained on annotations of World War 1-era Chronicling America pages, including annotations made by volunteers as part of the&nbsp;<a href="https://labs.loc.gov/work/experiments/beyond-words/">Beyond Words</a>&nbsp;crowdsourcing project.</p> <p>source:<a href="https://news-navigator.labs.loc.gov/"> https://news-navigator.labs.loc.gov/</a></p> </blockquote> <p>One of these categories is &#39;photographs&#39;. This dataset contains a sample of these images with additional labels indicating if the photograph has one or more of the following labels: &quot;human&quot;, &quot;animal&quot;, &quot;human-structure&quot; and &quot;landscape&quot;</p> <p>The data is organised as follows:</p> <ul> <li>The images themselves can be found in `images.zip`</li> <li>`newspaper-navigator-sample-metadata.csv` contains metadata about each image drawn from the Newspaper Navigator Dataset.</li> <li>`multi_label.csv` contains the labels for the images as a CSV file</li> <li>`annotations.csv` conains the labels for the images with additional metadata</li> </ul> <p>This dataset was created for use in an under-review Programming Historian tutorial (<a href="http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt2">http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt2</a>) The primary aim of the data was to provide a realistic example dataset for teaching computer vision for working with digitised heritage material. The data is shared here since it may be useful for others. <strong>This data documentation is a work in progress and will be updated when the Programming Historian tutorial is released publicly. </strong></p> <p>The metadata CSV file contains the following columns:</p> <p>- filepath<br> - pub_date<br> - page_seq_num<br> - edition_seq_num<br> - batch<br> - lccn<br> - box<br> - score<br> - ocr<br> - place_of_publication<br> - geographic_coverage<br> - name<br> - publisher<br> - url<br> - page_url<br> - month<br> - year<br> - iiif_url</p>

openother-openJan 2022View details →
zenodo44/100

The ALPIN Sentiment Dictionary: Austrian Language Polarity in Newspapers

<p>These datasets are part of the submitted paper for the LREC2022 conference entitled: &quot;The ALPIN Sentiment Dictionary: Austrian Language Polarity in Newspapers&quot;</p> <p>The various data sources, as well as the methodology, are explained in detail in the research paper which will be available soon.</p> <p>ALPIN stands for Austrian Language Polarity in Newspapers. The dictionary consists of three different parts which were merged together:</p> <ul> <li>Austrian Media Corpus: AMC (AMC_v1.0.csv)</li> <li>STANDARD posts: STP (STP_v1.0.csv)</li> <li>Austriacisms: AUT (AUT_v1.0.csv)</li> </ul> <p>Austrian Media Corpus (AMC) (Ransmayr et al., 2017) &amp; STANDARD posts (STP) (Schabus et al., 2017) rely on the SPLM algorithm as used in SentiDraw (Sharma &amp; Dutta 2021). Austriacisms (AUT) was generated by using the Best-Worst scaling (BWS) (Kiritchenko and Mohammad, 2017b). The AUT list was collected from the &ldquo;Variantenw&ouml;rterbuch des Deutschen&rdquo; (Ammon et al., 2016) (thereby only selecting those words that only surface in Austrian German and in no other variety of German) and an austriacism list of Wikipedia (https://de.wikipedia.org/wiki/Liste_von_Austriazismen).</p> <p>The scores are scaled to the interval [-1, 1] using the min-max-abs scaling, ranging from negative to positive.</p> <p>References:<br> Sharma, S. S., &amp; Dutta, G. (2021). SentiDraw: Using star ratings of reviews to develop domain specific sentiment lexicon for polarity determination. Information Processing &amp; Management, 58(1), 102412.<br> Kiritchenko, S. and Mohammad, S. M. (2017b). Capturing reliable fine-grained sentiment associations by crowdsourcing and best-worst scaling.<br> Schabus, D., Skowron, M., &amp; Trapp, M. (2017). One Million Posts: A Data Set of German Online Discussions. Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, 1241&ndash;1244. https://doi.org/10.1145/3077136.3080711<br> Ransmayr, J., M&ouml;rth, K., &amp; Ďurčo, M. (2017). AMC (Austrian Media Corpus). In Korpusbasierte Forschungen zum &ouml;sterreichischen Deutsch. In Digitale Methoden der Korpusforschung in &Ouml;sterreich (= Ver&ouml;ffentlichungen zur Linguistik und Kommunikationsforschung Nr. 30) (pp. 27&ndash;38). Verlag der &Ouml;sterreichischen Akademie der Wissenschaften.<br> Ammon, U., Bickel, H., &amp; Ebner, J. (2016). Variantenw&ouml;rterbuch des Deutschen : die Standardsprache in &Ouml;sterreich, der Schweiz, Deutschland, Liechtenstein, Luxemburg, Ostbelgien und S&uuml;dtirol sowie Rum&auml;nien, Namibia und Mennonitensiedlungen. Walter de Gruyter.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Newspapers_09/29/22

Documentation material from the Mastic pilot of the Mingei project

opencc-by-sa-4.0Sep 2022View details →
zenodo44/100

KB newspaper models for ShiCo

<p>The word2vec models were created from the KB news paper archive collection (a collection of digitised newspaper, created by the National Library of the Netherlands). Each word2vec model is created using newspaper articles published within a 10 year time frame. These models were created using the Gensim word2vec implementation with default parameters: Skip-gram architecture, with hierarchical softmax (no negative sampling), vector dimensionality of 300, window size of 5, and minimum word frequency of 5. Models are saved in binary word2vec format.</p>

opencc-by-sa-4.0Mar 2018View details →
zenodo44/100

Named-Entity Recognition for Modern Tibetan Newspapers: Tagset, Guidelines and Training Data

<p>This dataset, tagset and guidelines were the output&nbsp;of a six-month incubator project on the feasibility of developing Named-Entity Recognition (NER) for modern Tibetan, primarily for use with contemporary Tibetan-language newspapers and media published inside the PRC.&nbsp;The project was carried out by the Mongolian and Inner Asian Studies Unit at Cambridge University&rsquo;s Department of Social Anthropology. It was funded by an incubator grant from Cambridge Language Sciences. The project title was&nbsp;&ldquo;Named-Entity Recognition in Tibetan and Mongolian Newspapers.&rdquo; The Project PI was Dr Hildegard Diemberger (Cambridge),&nbsp;the Coordinator and Lead Author was Dr Robert Barnett (SOAS), and Senior Advisers were Dr Nathan Hill (SOAS), Dr Marieke Meelen (Cambridge), and Dr Thomas White (Cambridge).&nbsp;<br> <br> Although some forms of NER and other NLP procedures have been developed within China for modern Tibetan (see Liu, Nuo <em>et al</em>, 2011), the data underlying those initiatives have not been made publicly available and their findings cannot be tested or reproduced. Significant work on developing NLP for Tibetan has been carried out outside China, but has focused largely on classical Tibetan and religious texts (see Hill &amp; Garrett, Edward, 2017).&nbsp;</p> <p>The Cambridge incubator project therefore produced a tagset, guidelines and training data for developing NER for modern Tibetan, with a focus on historical and political analysis of contemporary newspapers, media and other public documents in Tibetan. We compiled 3.11m syllables of data in Tibetan extracted from articles downloaded from Chinese-language news aggregator sites within China, primarily tibet.cpc.people.com.cn and tibet.people.com.cn. From this data, we selected texts containing 280,000 syllables in Tibetan, grouped in 26,000 utterances/sentences (available on request). Using Lighttag, an online annotation site, we developed a tagset for NER consisting of 17 tags (and one for wrong segmentation if using segmented data). We annotated approximately 186,000 syllables, leading to 9,884 annotations. Of these, after discounting flawed data, we produced training data containing c.6,700&nbsp;annotations.&nbsp; We carried out the secondary, manual review offline (for our method of converting Lighttag&nbsp;data for offline review, see the attached report &ldquo;Using Spreadsheets to Review Annotations Offline.pdf&rdquo;), and found an error rate of 3.6%. The final total of reviewed annotations was 6,624.&nbsp;</p> <p>The dataset, tagset, guidelines and reports were developed and documented by Robert Barnett, with assistance from Tsering Samdrup, Dr Hill and Dr Meelen. Primary annotation was by Tsering Samdrup, assisted by Dr Barnett.<br> <br> The datasets published here include:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p> <ol> <li>The <strong>tagseet guidelines and annotation manual</strong>, including the 17-tag tagset, guidelines, and recommendations&nbsp;(&quot;NER for Modern Tibetan-tagset and guidelines.pdf&quot;).</li> <li>The <strong>tagged training data </strong>in .csv format (&quot;Tibetan NER Training Data-tagged, reviewed wth context-v10-UTF-8.csv&quot;) and .xls format (&quot;Tibetan NER Training Data-tagged with context-v10-UTF-8.xlsx&quot;). This includes&nbsp;6,624&nbsp;reveiwed annotations, arranged according to the Tibetan alphabet&nbsp;together&nbsp;with the tags and context (utterance) for each annotation.</li> <li>The <strong>raw annotation results </strong>downloaded&nbsp;from Lighttag as .json files&nbsp;(&quot;Raw Training Data for NER in Modern Tibetan -Jobs2-11-JSON.zip&quot;) and as .xls files&nbsp;(&quot;Training Data for NER in Modern Tibetan -Jobs2-11-XLS.zip&quot;). These include&nbsp;10 &quot;tasks&quot; or datasets of articles scraped from Tibetan-language websites within Tibet.&nbsp;&nbsp;&nbsp;&nbsp;</li> <li>A <strong>guide to preparing Lighttag annotation results for manual review offline </strong>(&ldquo;Using Spreadsheets to Review Annotations Offline.pdf&rdquo;).</li> </ol> <p>The project&#39;s findings regarding the status of NER and NLP for vertical Mongolian are available at DOI: 10.5281/zenodo.5103499.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Diachronic and diatopic word embeddings from newspapers digitised by the British Library (1830-1889): North and South England

<p>Diachronic&nbsp;word embeddings (decade-level) trained with Word2Vec (via Gensim) on different geographic subcorpora of the Heritage Made Digital British and the Living with Machines&nbsp;historical newspaper collections:</p> <p>- North England (north.zip)</p> <p>- South England (south.zip)</p> <p>At the moment, for each subcorpus, Word2Vec models are available for each decade in the period 1830-1889. More models are on the way for the following:</p> <p>- each decade in the periods 1780-1829 and 1890-1920 for both North and South England.</p> <p>- diachronic models for the following regions: Scotland, Wales, and Midlands.</p> <p>The models were trained using the following parameters:</p> <pre><code>sg = True min_count = 1 window = 5 vector_size = 200 epochs = 5</code></pre> <p>Like the embeddings in&nbsp;<a href="https://doi.org/10.5281/zenodo.7181682">this repository</a>, the model for each decade was aligned to the most recent one with Orthogonal Procrustes.</p> <p>See related GitHub repository for the full documentation:&nbsp;<a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a>.</p> <p>Project website (Living with Machines):&nbsp;<a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p> <p>Data related to:&nbsp;Nilo Pedrazzini &amp; Barbara McGillivray,&nbsp;<em>Diachronic and diatopic word embeddings from British historical newspapers</em>, presented at&nbsp;AIUCD (Convegno dell&rsquo;Associazione per l&rsquo;Informatica Umanistica e la Cultura Digitale) in Siena (Italy), June 2023.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Argentina news from Pagina12 Newspaper

<p>This dataset contains information from the web page of the Argentinian newspaper &quot;Pagina12&quot; (https://www.pagina12.com.ar/). The dataset was developed with a Python code in the context of academical studies (Data Science master). The final purpose was only for academic research.</p> <p>The dataset includes different fields of an article link of the cited newspaper, specifically the code was implemented in Python using Srapy library. Hence, the code scrap the final link of each notice or new from the different sections. For each notice or new, the following fields have been extracted: title, date, author, section and url of the new.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper

<p>Corpora used in the publication:</p> <ul> <li>Cristina Espa&ntilde;a-Bonet. 2023. <strong>Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper. </strong>In <em>Findings of the Association for Computational Linguistics: EMNLP 2023</em>, Singapore. Pages 11757&ndash;11777. Association for Computational Linguistics.</li> </ul> <p>Three corpora are included:</p> <ol> <li>Newspaper articles extracted from the OSCAR corpus in English, German, Spanish and Catalan automatically annotated for political stance (left vs right) and topic</li> <li>Newspaper-like article generations by different versions of ChatGPT for 101 topics in the 4 languages</li> <li>Newspaper-like article generations by Bard for 101 topics in the 4 languages</li> </ol> <p>See the README file and the original article for further details.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

NewsEye / READ OCR training dataset from French Newspapers (18th, 19th, early 20th C.)

<p>The dataset comprises French newspaper pages from 18th, 19th and early 20th century with carefully corrected text. The page images were provided by the&nbsp;<a href="https://www.bnf.fr/en">French National Library</a> and comprise 127 pages (training set) and 8 pages (validation set). The data are formed according to the PAGE format (cf.&nbsp;Cf.&nbsp;<a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a>&nbsp;and the&nbsp;<a href="http://read.transkribus.eu/">READ </a>project.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

BioPropaPhenKG on Online Newspapers and Medical Articles

<p><br>The coronavirus disease (COVID-19) spread rampantly around the world at the beginning of 2020 before the governments of each country could prevent it by making decisions based on medical data analysis. With proper formalization, the terabytes of new textual data available online every day could have been used for the early description and detection of cases of this virus. Since then, the number of Event-Based Surveillance (EBS) applications has increased exponentially. These applications aim to mine channels of unstructured information to detect signs of possible public health events' progression. However, one problem with such systems is the need for expert intervention to define which event will be captured, which relevant terms should be used in the search, and to analyze the events to modify the search procedure constantly. Another problem is that many of these applications do not consider both spatial and temporal characteristics. Addressing such limitations, this datasets presents a novel approach. We propose the use of BioPropaPhenKG to replace such systems. In this dataset, BioPropaPhen was enhanced with information comming from unstructured texts from online newspapers and medical articles. BioPropaPhenKG, its ontology and other useful information can be found in <a href="../records/10911980">https://zenodo.org/records/10911980</a>. The code used for this use case can be found in <a href="https://github.com/Gabriel382/DDPF-Health-Risks">https://github.com/Gabriel382/DDPF-Health-Risks</a> . Finally, the datasets used where UMLS MetamorphoSys, OpenStreetMaps, Wikidata, <a href="https://aylien.com/blog/free-coronavirus-news-dataset">Aylien</a> (data only from November of 2019) and&nbsp;<a href="https://allenai.org/data/cord-19">CORD-19</a> (data only from December of 2019).&nbsp;</p> <p>&nbsp;</p> <p>To read, you just need to load it with Neo4j:4.4.3. Alternatively, you can open it with docker using the following command:&nbsp;</p> <p>docker run --interactive --tty --rm \<br>&nbsp; &nbsp; --publish=7474:7474 --publish=7687:7687 \<br>&nbsp; &nbsp; --volume=/path-to-data-folder:/data --user="$(id -u):$(id -g)"\<br>&nbsp; &nbsp; neo4j:4.4.3 \<br>neo4j-admin load --from=/data/BioPropaPhenKG-Journal-Medical.dump --database "neo4j" --force</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

BioPropaPhenKG on Multi-Relation Extraction Methods on Online Newspapers

<p>The coronavirus disease (COVID-19) spread rampantly around the world at the beginning of 2020 before the governments of each country could prevent it by making decisions based on medical data analysis. With proper formalization, the terabytes of new textual data available online every day could have been used for the early description and detection of cases of this virus. Since then, the number of Event-Based Surveillance (EBS) applications has increased exponentially. These applications aim to mine channels of unstructured data to detect signs of possible public health events. However, one problem with such systems is the need for expert intervention to define which event will be captured, which relevant terms should be used in the search, and to analyze the events to modify the search procedure constantly. Another problem is that many of these applications do not consider both spatial and temporal characteristics. Addressing such limitations, this article presents a novel approach. We propose the biomedical domain specialization of the Core Propagation Phenomenon Ontology (PropaPhen) to capture spatiotemporal characteristics of the propagation of health-related phenomena. We also propose the Description-Detection-Framework (DDF), which leverages PropaPhen, UMLS, and OpenStreetMaps to detect new medical events automatically. Finally, we demonstrate a use case with experiments on extracts from online newspapers about COVID-19. The results show that DDF can be useful for detecting clusters of suspicious cases of possible emerging health-related phenomena.</p> <p>BioPropaPhenKG, its ontology and other useful information can be found in&nbsp;<a href="../records/10911980">https://zenodo.org/records/10911980</a>. The code used for this use case can be found in <a href="https://github.com/Gabriel382/DDPF-Health-Risks">https://github.com/Gabriel382/DDPF-Health-Risks</a> . Finally, the datasets used where UMLS MetamorphoSys, OpenStreetMaps, Wikidata, <a href="https://aylien.com/blog/free-coronavirus-news-dataset">Aylien</a> (data only from November of 2019).</p> <p>&nbsp;</p> <p>To read, you just need to load it with Neo4j:4.4.3. Alternatively, you can open it with docker using the following command:&nbsp;</p> <p>docker run --interactive --tty --rm \<br>&nbsp; &nbsp; --publish=7474:7474 --publish=7687:7687 \<br>&nbsp; &nbsp; --volume=/path-to-data-folder:/data --user="$(id -u):$(id -g)"\<br>&nbsp; &nbsp; neo4j:4.4.3 \<br>neo4j-admin load --from=/data/BioPropaPhenKG-Journal-MultiRE.dump --database "neo4j" --force</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Dataset for Logical-layout analysis on French historical newspapers

<p><strong>Dataset for Logical-layout analysis on French historical newspapers</strong></p> <p>This dataset is intended for training and testing Logical Layout Analysis and recognition system on French historical documents published between 1900 and 1950. The original data is part of the &quot;<a href="https://gallica.bnf.fr/services/engine/search/sru?operation=searchRetrieve&amp;exactSearch=false&amp;version=1.2&amp;query=%28colnum%20adj%20%22Appartient%20%C3%A0%20l%27ensemble%20documentaire%20:%20FrancComt1%22%29">Fond r&eacute;gional: Franche-Comt&eacute;</a>&quot;, which is curated by <a href="https://gallica.bnf.fr/accueil/fr/content/accueil-fr?mode=desktop">Gallica</a>, the digital portal of the Biblioth&egrave;que Nationale de France (BnF). This dataset has the following structure:</p> <p>├── train<br> &nbsp; ├── 1c<br> &nbsp;&nbsp;&nbsp; ├── cb32836282t<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── cb32836282t.xml<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── bpt6k112325g<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── bpt6k112325g.xml<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── truelabels_block.csv<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── truelabels_line.csv<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ├── &hellip;<br> &nbsp;&nbsp;&nbsp; ├── &hellip;<br> &nbsp; ├── 2c<br> &nbsp; ├── 3c+<br> └── test<br> &nbsp; ├── 1c<br> &nbsp; ├── 2c<br> &nbsp; └── 3c+</p> <p>The dataset is divided into a train and a test set. The train and test datasets have been designed to cover as much as possible the various possible layouts that exist in the &quot;Fond r&eacute;gional: Franche-Comt&eacute;&quot; dataset. To do so, we have divided them into three layout types:<br> &nbsp; &bull; <strong>1c</strong>: documents where the text is displayed in one column, as in books;<br> &nbsp; &bull; <strong>2c</strong>: documents where the text is displayed into two columns;<br> &nbsp; &bull; <strong>3c+</strong>: documents where there are at least 3 columns of text, as in newspapers.</p> <p>Each of the 1c, 2c, and 3c+ folder contains subfolders prefixed by &lsquo;cb&rsquo;, which contain a collection of documents. For instance, &laquo; cb32836282t &raquo; is the identifier used in Gallica for &laquo; Le Petit &eacute;cho du 21e R&eacute;giment d&#39;infanterie &raquo;, a French military periodical published during WWI. An XML file with the same name, for instance &laquo;cb32836282t.xml &raquo;, contains metadata about the collection, such as its title, publisher, creator, number of issues, etc. This XML file serves only to describe the collection, and is not to be used for Logical-Layout analysis.</p> <p>The issues in each collection can be found in the subfolders prefixed with &laquo; bpt &raquo;. For instance, &laquo; bpt6k112325g &raquo; is the identifier used in Gallica for an issue published in September 1917 of &laquo; Le Petit &eacute;cho du 21e R&eacute;giment d&#39;infanterie &raquo;. The information about each issue is given in three files, which are described below:</p> <p><strong>1-bptXXXXXXXXXX.xml </strong><br> The original data, as collected from Gallica. The most important tags of this document and their values are described below:<br> &nbsp; &bull; <strong>oai</strong>: metadata about the document, such as its author, title, publisher, original publication date, number of issues, &hellip;<br> &nbsp; &bull; <strong>image_url</strong>: the url to the document&rsquo;s scan (in high resolution)<br> &nbsp; &bull; <strong>pagination</strong>: a description of each page in the document (size of the page, if it contains a table of content or not, &hellip;)<br> &nbsp; &bull; <strong>num_pages</strong>: the total number of pages in the document<br> &nbsp;&nbsp;&bull; <strong>ocr</strong>: the OCR representation of the document in the XML ALTO format</p> <p>The XML ALTO format provides the text content and physical layout of documents in the following manner. Lines of text are contained in TextLine tags, which in their turn contain String tags for words and SP tags for spaces. TextLine tags are grouped into blocks in TextBlock tags. Sometimes, TextBlock tags are also grouped into ComposedBlock tags. TextBlock and TextLine tags have the following attributes:<br> &nbsp; &bull; <strong>Id</strong> : the tag&rsquo;s identifier<br> &nbsp; &bull; <strong>Height</strong>, <strong>Width</strong> : the text height and width<br> &nbsp; &bull; <strong>Vpos</strong> : the vertical position of the text on the page. The higher the value, the lower the word is on the page<br> &nbsp; &bull; <strong>Hpos</strong> : the horizontal position of the text on the page. The higher the value, the further on the right the text is on the page<br> &nbsp; &bull; <strong>Language</strong> : the language of the text (only for TextBlock tags).</p> <p>Among the attributes listed above, some TextBlock tags also have a Type attribute. This attribute contains logical labels of the lines in the block. In this dataset it appears most often for tables or advertisements. Overall, TextBlock tags that have a Type attribute are rare in this dataset (about 4 % only).</p> <p><strong>Note</strong>: The original scan of every document is accessible on the Gallica website, using the URL https://gallica.bnf.fr/ark:/12148/&lt;IDENTIFIER&gt;, where &lt;IDENTIFIER&gt; should be replaced by the id of the document (e.g.: bpt6k112325g) or the collection (e.g.: cb32836282t).</p> <p><strong>2-truelabels_block.csv </strong><br> A CSV file where each line corresponds to a TextBlock tag from the file bptXXXXXXXXXX.xml. This CSV file contains the following columns:<br> &nbsp; &bull; <strong>page</strong>: the page on which the TextBlock tag is located<br> &nbsp; &bull; <strong>block_id</strong>: the id of the TextBlock tag<br> &nbsp; &bull; <strong>first_last_line</strong>: the text content of the first and last TextLine tags inside this TextBlock tag<br> &nbsp; &bull; <strong>classes</strong>: the logical label(s) associated with this TextBlock tag</p> <p>The possible values in the column classes are : Text, Title, Header and Other.</p> <p><strong>3-truelabels_line.csv </strong><br> A CSV file where each line corresponds to a TextLine tag from the file bptXXXXXXXXXX.xml. This CSV file contains the following columns:<br> &nbsp; &bull; <strong>page</strong>: the page where the TextLine tag is located<br> &nbsp; &bull; <strong>block_id</strong>: the id of the TextBlock tag that contains this TextLine tag<br> &nbsp; &bull; <strong>line_id</strong>: the id of this TextLine tag<br> &nbsp; &bull; <strong>text_line</strong>: the text content of this TextLine tag<br> &nbsp; &bull; <strong>classes</strong>: the logical label(s) associated with this TextLine tag</p> <p>The possible values in the column classes are : Text, Firstline, Title, Header and Other. Firstline indicates the &laquo; first line &raquo; of a paragraph.<br> &nbsp;</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record