Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

35

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

35 results for “Digital Library”

Learn how ShareScore rates datasets ↗
zenodo52/100

13/1 Sferamundi di Grecia. Prima parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/1 Sferamundi di Grecia. Prima parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

13/2 Sferamundi di Grecia. Seconda parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/2 Sferamundi di Grecia. Seconda parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

13/6 Sferamundi di Grecia. Sesta parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/6 Sferamundi di Grecia. Sesta parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

13/5 Sferamundi di Grecia. Quinta parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/5 Sferamundi di Grecia. Quinta parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

13/4 Sferamundi di Grecia. Quarta parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/4 Sferamundi di Grecia. Quarta parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo52/100

13/3 Sferamundi di Grecia. Terza parte - Progetto Mambrino Digital Library

<p>Dataset of the digital scholarly edition of the Italian book of chivalry&nbsp;<em>13/3 Sferamundi di Grecia. Terza parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>

opencc-by-sa-4.0May 2024View details →
zenodo48/100

Berlin State Library (2024). Metadata of the Digitized Collections of the Berlin State Library (SBB)

<p>The motivation for creating this dataset was to enable research on the basis of metadata which are available in a cultural heritage institution on a large scale. Libraries such as the Staatsbibliothek zu Berlin &ndash; Berlin State Library (SBB) typically provide three kinds of data: Images (scans of books, illustrations contained in the scanned material, or else), texts (OCR'd from digitized books or manuscripts), and metadata. However, metadata form an underresearched resource, which is lamentable: These metadata are of a high quality since they have been established by trained librarians, archivists, or other cultural heritage practitioners. The publication of a set of metadata of more than 200.000 works aims therefore at providing an underresearched high-quality type of data. The basic interest of the funder in this data publication is the stimulation of innovation.</p> <p>The dataset consists of a single table containing the metadata of all 219.419 works which were available in the Digitized Collections of the Berlin State Library (SBB) on July 29th, 2024. The size of the .parquet file is about 46 MB.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Metadata, Title Pages, and Network Graph of the Digitized Content of the Berlin State Library (146,000 items)

<p>The data set has been downloaded via the OAI-PMH endpoint of the Berlin State Library/Staatsbibliothek zu Berlin&rsquo;s Digitized Collections (<a href="https://digital.staatsbibliothek-berlin.de/oai">https://digital.staatsbibliothek-berlin.de/oai</a>) on March 1<sup>st</sup> 2019 and converted into common tabular formats on the basis of the provided Dublin Core metadata. It contains 146,000 records.</p> <p>In addition to the bibliographic metadata, representative images of the works have been downloaded, resized to a 512 pixel maximum thumbnail image and saved in JPEG format. The image data is split into title pages and first pages. Title pages have been derived from structural metadata created by scan operators and librarians. If this information was not available, first pages of the media have been downloaded. In case of multi-volume media, title pages are not available.</p> <p>In total, 141,206 images title/first pages are available.</p> <p>&nbsp;</p> <p>Furthermore, the tabular data has been cleaned and extended with geo-spatial coordinates provided by the OpenStreetMap project (<a href="https://www.openstreetmap.org">https://www.openstreetmap.org</a>). The actual data processing steps are summarized in the next section. For the sake of transparency and reproducibility, the original data taken from the OAI-PMH endpoint is still present in the table.</p> <p>&nbsp;</p> <p>To conclude with, various graphs in GML file format are available that can be loaded directly into graph analysis tools such as Gephi (<a href="https://gephi.org/">https://gephi.org/</a>).</p> <p>&nbsp;</p> <p>The implementation of the data processing steps (incl. graph creation) are available as a Jupyter notebook provided at <a href="https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb">https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb</a>.</p> <p>&nbsp;</p> <p>Tabular Metadata</p> <p>&nbsp;</p> <p>The metadata is available in Excel (cleanedData.xlsx) and CSV (cleanedData.csv) file formats with equal content.</p> <p>The table contains the following columns. Italique columns have not been processed.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>title</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The title of the medium</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>creator</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Its creator (family name, first name)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>subject</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A collection&rsquo;s name as provided by the library</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>type</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The type of medium</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>format</em> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A MIME type for full metadata download</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>identifier</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; An additional identifier (most often the PPN)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>language</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A 3-letter language code of the medium</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>date</em>&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The date of creation/publication or a time span</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>relation</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A relation to a project or collection a medium has been digitized for.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>coverage</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The location of publication or origin (ranging from cities to continents)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>publisher</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The publisher of the medium.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>rights</em>&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Copyright information.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>PPN</em>&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The unique identifier that can be used to find more information about the current medium in all information systems of Berlin State Library/Staatsbibliothek zu Berlin.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; spatialClean&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; In case of multiple entries in coverage, only the first place of origin has been extracted. Additionally, characters such as question marks, brackets, or the like have been removed. The entries have been normalized regarding whitespaces and writing variants with the help of regular expressions.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; dateClean&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; As the original date may contain various format variants to indicate unclear creation dates (e.g., time spans or question marks), this field contains a mapping to a certain point in time.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; spatialCluster &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The cluster ID determined with the help of the Jaro-Winkler distance on the spatialClean string. This step is needed because the spatialClean fields still contain a huge amount of orthographic variants and latinizations of geographic names.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; spatialClusterName&nbsp;&nbsp; A verbal cluster name (controlled manually).</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; latitude&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The latitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; longitude&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The longitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; century&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A century derived from the date.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; textCluster&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A text cluster ID on the basis of a k-means clustering relying on the title field with a vocabulary size of 125,000 using the tf*idf model and k=5,000.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; creatorCluster &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A text cluster ID based on the creator field with k=20,000.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; titleImage&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The path to the first/title page relative to the img/ subdirectory or None in case of a multi-volume work.</p> <p>Other Data</p> <p>&nbsp;</p> <p><em>graphs.zip</em></p> <p>&nbsp;</p> <p>Various pre-computed graphs.</p> <p><em>&nbsp;</em></p> <p><em>img.zip</em></p> <p>&nbsp;</p> <p>First and title pages in JPEG format.</p> <p>&nbsp;</p> <p><em>json.zip</em></p> <p>&nbsp;</p> <p>JSON files for each record in the following format:</p> <p>&nbsp;</p> <p>ppn&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;PPN57346250X&quot;</p> <p>dateClean&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;1625&quot;</p> <p>title&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;M. Georgii Gutkii, Gymnasii Berlinensis Rectoris Habitus Primorum Principiorum, Seu Intelligentia; Annexae Sunt Appendicis loco Disputationes super eodem habitu tum in Academia Wittebergensi, tum in Gymnasio Berlinensi ventilatae&quot;</p> <p>creator&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Gutke, Georg&quot;</p> <p>spatialClusterName&nbsp;&nbsp; &quot;Berlin&quot;</p> <p>spatialClean&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Berolini&quot;</p> <p>spatialRaw&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Berolini&quot;</p> <p>mediatype&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;monograph&quot;</p> <p>subject&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Historische Drucke&quot;</p> <p>publisher&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Kallius&quot;</p> <p>lat&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;52.5170365&quot;</p> <p>lng&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;13.3888599&quot;</p> <p>textCluster&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;45&quot;</p> <p>creatorCluster&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;5040&quot;</p> <p>titleImage&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;titlepages/PPN57346250X.jpg&quot;</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

Extracted Illustrations of the Berlin State Library's Digitized Collections (part 1 of 4)

<p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library&#39;s Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><em>Change Log for Version</em></p> <ol> <li>original dataset</li> <li>added color histograms (RGB, separated by channel) in JSON and Python pickle format as extracted by the <a href="https://pillow.readthedocs.io/en/stable/">Pillow</a> package (see https://github.com/elektrobohemian/StabiHacks/tree/master/image-tools)</li> </ol> <p><strong>This is part 1 of 4. The following datasets contain the other ZIP files (8 files in total):</strong></p> <ul> <li><a href="https://doi.org/10.5281/zenodo.2598145">https://doi.org/10.5281/zenodo.2598145</a></li> <li><a href="https://doi.org/10.5281/zenodo.2598261">https://doi.org/10.5281/zenodo.2598261</a></li> <li><a href="https://doi.org/10.5281/zenodo.2598270">https://doi.org/10.5281/zenodo.2598270</a></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

OCR fulltexts of the Digital Collections of the Berlin State Library (DC-SBB)

<p>The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945.</p> <p>At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.<br> For each page with OCR text, the language has been determined by <em>langid </em>(Lui/Baldwin 2012).</p> <p>corpus-entropy.pkl &nbsp; &nbsp;&nbsp; entropy rate per document page</p> <p>corpus-language.pkl&nbsp;&nbsp; language per document page</p> <p>corpus.zip &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; fulltext corpus (extracts to .txt format)</p> <p>de_corpus.zip &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; German sub-corpus (extracts to .txt format)</p> <p>selection_de.pkl&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Selection list of German documents</p> <p>xml2csv_alto.csv&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; fulltext corpus per document page (incl.OCR word confidences)</p> <p>&nbsp;</p> <p><em>Sources</em></p> <p>Marco Lui and Timothy Baldwin. 2012. Langid.py:</p> <p>An off-the-shelf language identification tool. In Proceedings of the ACL 2012 System Demonstrations,</p> <p>ACL &rsquo;12, pages 25&ndash;30, Stroudsburg, PA, USA. Association for Computational Linguistics</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

ACM-Digital Library (ACM)

<p>ACM-Digital Library (ACM): a subset of the ACM Digital Library with 24, 897 documents containing articles related to Computer Science. We considered only the first level of the taxonomy adopted by ACM, where each document is assigned to one of 11 classes.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_&lt;k&gt;.pkl:&nbsp;&nbsp;pandas DataFrame with k-cross validation partition.</p> <p>The .zip contains all aforementioned files + the tfidf representation in the CSR matrix format.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Mnemosine Digital Library Functionality

<p>Source:&nbsp;</p> <p><em><span><span>&nbsp;</span>Mnemosine </span></em><span>and <em>Clavy</em>: applications for the development of specific libraries for teaching and research. Office for the Transfer of Research Results (</span><span><a href="https://www.ucm.es/otri/complutransfer-mnemosine-y-clavy-aplicaciones-para-la-gestacion-de-bibliotecas-especificas-para-la-docencia-y-la-investigacion"><span>OTRI</span></a></span><span>) <a href="https://www.ucm.es/otri/complutransfer-mnemosine-y-clavy-aplicaciones-para-la-gestacion-de-bibliotecas-especificas-para-la-docencia-y-la-investigacion">https://www.ucm.es/otri/complutransfer-mnemosine-y-clavy-aplicaciones-para-la-gestacion-de-bibliotecas-especificas-para-la-docencia-y-la-investigacion</a></span></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Digital Scholarship services and supports - an overview from Irish Research and National Libraries - Data with Comments

<p>This dataset contains survey data with comments (cleaned, direct references to institutions removed) and broken down by survey&nbsp;question sections. Some of the data has been converted to counts where it was used to generate the charts.</p> <p>The data is the output of a 2018 CONUL (Ireland&rsquo;s consortium of research and national libraries) survey focusing on Digital Scholarship services and supports from a Irish Research and National Library&nbsp;perspective.</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Extracted Illustrations of the Berlin State Library's Digitized Collections (part 4 of 4)

<p><strong>This is part 4 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre> <p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library&#39;s Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><strong>This is part 4 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Extracted Illustrations of the Berlin State Library's Digitized Collections (part 3 of 4)

<p><strong>This is part 3 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre> <p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library&#39;s Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><strong>This is part 3 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Extracted Illustrations of the Berlin State Library's Digitized Collections (part 2 of 4)

<p><strong>This is part 2 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre> <p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library&#39;s Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><strong>This is part 2 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre>

opencc-by-4.0Mar 2019View details →
zenodo40/100

Berlin State Library (2023). Fulltexts of the Digitized Collections of the Berlin State Library (SBB)

<p>The motivation for creating this dataset was to enable research on the basis of fulltexts which are available in a cultural heritage institution on a large scale. Libraries such as the Berlin State Library (SBB) typically provide three kinds of data: Images (scans of books, illustrations contained in the scanned material, or else), metadata (descriptive data providing information on the digitized item) and texts. The latter are usually received via an implementation of optical character recognition (OCR) of the digitized books or manuscripts. In the <a href="https://digital.staatsbibliothek-berlin.de/">digitized collections of the Staatsbibliothek zu Berlin (SBB)</a>, the fulltexts can be downloaded manually, item by item. The publication of a set of about 5 million OCR&rsquo;d pages alleviates the accessibility of the fulltexts. The basic funding interest in this data publication is the stimulation of innovation.</p> <p>The dataset consists of a single sqlite database containing ALL the fulltexts available in the digitized collections of the Berlin State Library as of August 21st, 2019, with information on languages and entropy added on March 1st, 2023. The size of the sqlite file is about 15.8 GB.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

linguistlist/multitree: MultiTree: A digital library of language relationships

No description provided.

opencc-by-sa-4.0Oct 2023View details →
zenodo36/100

A comparative usability analysis of eye-tracking and mouse click data taken from digital libraries

<p>This dataset is the result of a study, in which we analyzed parallels and differences between clicks as well as eye movements on two different digital library homepages. For this analysis we used diverse tracking tools for mouse clicks and eye tracking data that where further studied with respect to specific areas of interest (AOI). &nbsp;</p> <p>The dataset contains two screenshots indicating the areas of interest (AOIs; entitled &ldquo;AreasOfInterest_Kartenportal.jpg and AreasOfInterest_Webportal.jpg), which separate the homepages into analyzable parts. It also contains eight screenshots of the homepages containing the total amount of collected clicks (each name starting with &ldquo;clicks&rdquo;) and two screenshots with the eye tracking heat maps (starting with &ldquo;Heatmap&rdquo;). The screenshots have directly been extracted from the click and eye tracking tools and matched with the before mentioned AOIs in order to gain the total count of clicks and views as well as the view duration the concerned area.</p> <p>All data are synthesized in a document containing three sheets with different tables: a first one with the initial data compilation for all AOIs of the two analyzed homepages (entitled &ldquo;Data&rdquo;), a second one with a more visual compiled data analysis for both homepages and all AOIs (entitled &ldquo;Data2) and last one with the duration of the view as well as the duration of the fixation and the compiled click data (entitled &laquo;&nbsp;Eye tracking study data&nbsp;&raquo;).</p>

opencc-zeroAug 2014View details →
zenodo36/100

Mnemosine Digital Library Architecture

<p>Mnemosyne digital library architecture, identifying relationships with external systems.</p>

opencc-by-4.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record