Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
35
datasets available to search
ShareScore release 0.7.1
Dataset results
35 results for “Digital Library”
13/1 Sferamundi di Grecia. Prima parte - Progetto Mambrino Digital Library
<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/1 Sferamundi di Grecia. Prima parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>
13/2 Sferamundi di Grecia. Seconda parte - Progetto Mambrino Digital Library
<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/2 Sferamundi di Grecia. Seconda parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>
13/6 Sferamundi di Grecia. Sesta parte - Progetto Mambrino Digital Library
<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/6 Sferamundi di Grecia. Sesta parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>
13/5 Sferamundi di Grecia. Quinta parte - Progetto Mambrino Digital Library
<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/5 Sferamundi di Grecia. Quinta parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>
13/4 Sferamundi di Grecia. Quarta parte - Progetto Mambrino Digital Library
<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/4 Sferamundi di Grecia. Quarta parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>
13/3 Sferamundi di Grecia. Terza parte - Progetto Mambrino Digital Library
<p>Dataset of the digital scholarly edition of the Italian book of chivalry <em>13/3 Sferamundi di Grecia. Terza parte</em>.</p> <p>It contains:</p> <ul> <li>transcription and commentary XML-TEI files (source.xml and commentary.xml)</li> <li>the eBook (in multiple formats)</li> <li>plain text file for computational anaysis</li> </ul> <p>The edition is part of the Progetto Mambrino Digital Library and has been developed within the PRIN 2017 Mapping Chivalry (Prot. 2017JA5XAR), in the context of the Project of Excellence "Inclusive Humanities" (2023-2027) of the Department of Foreign Languages and Literatures of the University of Verona.</p>
Berlin State Library (2024). Metadata of the Digitized Collections of the Berlin State Library (SBB)
<p>The motivation for creating this dataset was to enable research on the basis of metadata which are available in a cultural heritage institution on a large scale. Libraries such as the Staatsbibliothek zu Berlin – Berlin State Library (SBB) typically provide three kinds of data: Images (scans of books, illustrations contained in the scanned material, or else), texts (OCR'd from digitized books or manuscripts), and metadata. However, metadata form an underresearched resource, which is lamentable: These metadata are of a high quality since they have been established by trained librarians, archivists, or other cultural heritage practitioners. The publication of a set of metadata of more than 200.000 works aims therefore at providing an underresearched high-quality type of data. The basic interest of the funder in this data publication is the stimulation of innovation.</p> <p>The dataset consists of a single table containing the metadata of all 219.419 works which were available in the Digitized Collections of the Berlin State Library (SBB) on July 29th, 2024. The size of the .parquet file is about 46 MB.</p>
Metadata, Title Pages, and Network Graph of the Digitized Content of the Berlin State Library (146,000 items)
<p>The data set has been downloaded via the OAI-PMH endpoint of the Berlin State Library/Staatsbibliothek zu Berlin’s Digitized Collections (<a href="https://digital.staatsbibliothek-berlin.de/oai">https://digital.staatsbibliothek-berlin.de/oai</a>) on March 1<sup>st</sup> 2019 and converted into common tabular formats on the basis of the provided Dublin Core metadata. It contains 146,000 records.</p> <p>In addition to the bibliographic metadata, representative images of the works have been downloaded, resized to a 512 pixel maximum thumbnail image and saved in JPEG format. The image data is split into title pages and first pages. Title pages have been derived from structural metadata created by scan operators and librarians. If this information was not available, first pages of the media have been downloaded. In case of multi-volume media, title pages are not available.</p> <p>In total, 141,206 images title/first pages are available.</p> <p> </p> <p>Furthermore, the tabular data has been cleaned and extended with geo-spatial coordinates provided by the OpenStreetMap project (<a href="https://www.openstreetmap.org">https://www.openstreetmap.org</a>). The actual data processing steps are summarized in the next section. For the sake of transparency and reproducibility, the original data taken from the OAI-PMH endpoint is still present in the table.</p> <p> </p> <p>To conclude with, various graphs in GML file format are available that can be loaded directly into graph analysis tools such as Gephi (<a href="https://gephi.org/">https://gephi.org/</a>).</p> <p> </p> <p>The implementation of the data processing steps (incl. graph creation) are available as a Jupyter notebook provided at <a href="https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb">https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb</a>.</p> <p> </p> <p>Tabular Metadata</p> <p> </p> <p>The metadata is available in Excel (cleanedData.xlsx) and CSV (cleanedData.csv) file formats with equal content.</p> <p>The table contains the following columns. Italique columns have not been processed.</p> <p>· <em>title</em> The title of the medium</p> <p>· <em>creator</em> Its creator (family name, first name)</p> <p>· <em>subject</em> A collection’s name as provided by the library</p> <p>· <em>type</em> The type of medium</p> <p>· <em>format</em> A MIME type for full metadata download</p> <p>· <em>identifier</em> An additional identifier (most often the PPN)</p> <p>· <em>language</em> A 3-letter language code of the medium</p> <p>· <em>date</em> The date of creation/publication or a time span</p> <p>· <em>relation</em> A relation to a project or collection a medium has been digitized for.</p> <p>· <em>coverage</em> The location of publication or origin (ranging from cities to continents)</p> <p>· <em>publisher</em> The publisher of the medium.</p> <p>· <em>rights</em> Copyright information.</p> <p>· <em>PPN</em> The unique identifier that can be used to find more information about the current medium in all information systems of Berlin State Library/Staatsbibliothek zu Berlin.</p> <p>· spatialClean In case of multiple entries in coverage, only the first place of origin has been extracted. Additionally, characters such as question marks, brackets, or the like have been removed. The entries have been normalized regarding whitespaces and writing variants with the help of regular expressions.</p> <p>· dateClean As the original date may contain various format variants to indicate unclear creation dates (e.g., time spans or question marks), this field contains a mapping to a certain point in time.</p> <p>· spatialCluster The cluster ID determined with the help of the Jaro-Winkler distance on the spatialClean string. This step is needed because the spatialClean fields still contain a huge amount of orthographic variants and latinizations of geographic names.</p> <p>· spatialClusterName A verbal cluster name (controlled manually).</p> <p>· latitude The latitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>· longitude The longitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>· century A century derived from the date.</p> <p>· textCluster A text cluster ID on the basis of a k-means clustering relying on the title field with a vocabulary size of 125,000 using the tf*idf model and k=5,000.</p> <p>· creatorCluster A text cluster ID based on the creator field with k=20,000.</p> <p>· titleImage The path to the first/title page relative to the img/ subdirectory or None in case of a multi-volume work.</p> <p>Other Data</p> <p> </p> <p><em>graphs.zip</em></p> <p> </p> <p>Various pre-computed graphs.</p> <p><em> </em></p> <p><em>img.zip</em></p> <p> </p> <p>First and title pages in JPEG format.</p> <p> </p> <p><em>json.zip</em></p> <p> </p> <p>JSON files for each record in the following format:</p> <p> </p> <p>ppn "PPN57346250X"</p> <p>dateClean "1625"</p> <p>title "M. Georgii Gutkii, Gymnasii Berlinensis Rectoris Habitus Primorum Principiorum, Seu Intelligentia; Annexae Sunt Appendicis loco Disputationes super eodem habitu tum in Academia Wittebergensi, tum in Gymnasio Berlinensi ventilatae"</p> <p>creator "Gutke, Georg"</p> <p>spatialClusterName "Berlin"</p> <p>spatialClean "Berolini"</p> <p>spatialRaw "Berolini"</p> <p>mediatype "monograph"</p> <p>subject "Historische Drucke"</p> <p>publisher "Kallius"</p> <p>lat "52.5170365"</p> <p>lng "13.3888599"</p> <p>textCluster "45"</p> <p>creatorCluster "5040"</p> <p>titleImage "titlepages/PPN57346250X.jpg"</p>
Extracted Illustrations of the Berlin State Library's Digitized Collections (part 1 of 4)
<p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library's Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><em>Change Log for Version</em></p> <ol> <li>original dataset</li> <li>added color histograms (RGB, separated by channel) in JSON and Python pickle format as extracted by the <a href="https://pillow.readthedocs.io/en/stable/">Pillow</a> package (see https://github.com/elektrobohemian/StabiHacks/tree/master/image-tools)</li> </ol> <p><strong>This is part 1 of 4. The following datasets contain the other ZIP files (8 files in total):</strong></p> <ul> <li><a href="https://doi.org/10.5281/zenodo.2598145">https://doi.org/10.5281/zenodo.2598145</a></li> <li><a href="https://doi.org/10.5281/zenodo.2598261">https://doi.org/10.5281/zenodo.2598261</a></li> <li><a href="https://doi.org/10.5281/zenodo.2598270">https://doi.org/10.5281/zenodo.2598270</a></li> </ul> <p> </p>
OCR fulltexts of the Digital Collections of the Berlin State Library (DC-SBB)
<p>The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945.</p> <p>At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.<br> For each page with OCR text, the language has been determined by <em>langid </em>(Lui/Baldwin 2012).</p> <p>corpus-entropy.pkl entropy rate per document page</p> <p>corpus-language.pkl language per document page</p> <p>corpus.zip fulltext corpus (extracts to .txt format)</p> <p>de_corpus.zip German sub-corpus (extracts to .txt format)</p> <p>selection_de.pkl Selection list of German documents</p> <p>xml2csv_alto.csv fulltext corpus per document page (incl.OCR word confidences)</p> <p> </p> <p><em>Sources</em></p> <p>Marco Lui and Timothy Baldwin. 2012. Langid.py:</p> <p>An off-the-shelf language identification tool. In Proceedings of the ACL 2012 System Demonstrations,</p> <p>ACL ’12, pages 25–30, Stroudsburg, PA, USA. Association for Computational Linguistics</p>
ACM-Digital Library (ACM)
<p>ACM-Digital Library (ACM): a subset of the ACM Digital Library with 24, 897 documents containing articles related to Computer Science. We considered only the first level of the taxonomy adopted by ACM, where each document is assigned to one of 11 classes.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_<k>.pkl: pandas DataFrame with k-cross validation partition.</p> <p>The .zip contains all aforementioned files + the tfidf representation in the CSR matrix format.</p>
Mnemosine Digital Library Functionality
<p>Source: </p> <p><em><span><span> </span>Mnemosine </span></em><span>and <em>Clavy</em>: applications for the development of specific libraries for teaching and research. Office for the Transfer of Research Results (</span><span><a href="https://www.ucm.es/otri/complutransfer-mnemosine-y-clavy-aplicaciones-para-la-gestacion-de-bibliotecas-especificas-para-la-docencia-y-la-investigacion"><span>OTRI</span></a></span><span>) <a href="https://www.ucm.es/otri/complutransfer-mnemosine-y-clavy-aplicaciones-para-la-gestacion-de-bibliotecas-especificas-para-la-docencia-y-la-investigacion">https://www.ucm.es/otri/complutransfer-mnemosine-y-clavy-aplicaciones-para-la-gestacion-de-bibliotecas-especificas-para-la-docencia-y-la-investigacion</a></span></p>
Digital Scholarship services and supports - an overview from Irish Research and National Libraries - Data with Comments
<p>This dataset contains survey data with comments (cleaned, direct references to institutions removed) and broken down by survey question sections. Some of the data has been converted to counts where it was used to generate the charts.</p> <p>The data is the output of a 2018 CONUL (Ireland’s consortium of research and national libraries) survey focusing on Digital Scholarship services and supports from a Irish Research and National Library perspective.</p>
Extracted Illustrations of the Berlin State Library's Digitized Collections (part 4 of 4)
<p><strong>This is part 4 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre> <p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library's Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><strong>This is part 4 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre>
Extracted Illustrations of the Berlin State Library's Digitized Collections (part 3 of 4)
<p><strong>This is part 3 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre> <p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library's Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><strong>This is part 3 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre>
Extracted Illustrations of the Berlin State Library's Digitized Collections (part 2 of 4)
<p><strong>This is part 2 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre> <p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library's Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><strong>This is part 2 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre>
Berlin State Library (2023). Fulltexts of the Digitized Collections of the Berlin State Library (SBB)
<p>The motivation for creating this dataset was to enable research on the basis of fulltexts which are available in a cultural heritage institution on a large scale. Libraries such as the Berlin State Library (SBB) typically provide three kinds of data: Images (scans of books, illustrations contained in the scanned material, or else), metadata (descriptive data providing information on the digitized item) and texts. The latter are usually received via an implementation of optical character recognition (OCR) of the digitized books or manuscripts. In the <a href="https://digital.staatsbibliothek-berlin.de/">digitized collections of the Staatsbibliothek zu Berlin (SBB)</a>, the fulltexts can be downloaded manually, item by item. The publication of a set of about 5 million OCR’d pages alleviates the accessibility of the fulltexts. The basic funding interest in this data publication is the stimulation of innovation.</p> <p>The dataset consists of a single sqlite database containing ALL the fulltexts available in the digitized collections of the Berlin State Library as of August 21st, 2019, with information on languages and entropy added on March 1st, 2023. The size of the sqlite file is about 15.8 GB.</p>
linguistlist/multitree: MultiTree: A digital library of language relationships
No description provided.
A comparative usability analysis of eye-tracking and mouse click data taken from digital libraries
<p>This dataset is the result of a study, in which we analyzed parallels and differences between clicks as well as eye movements on two different digital library homepages. For this analysis we used diverse tracking tools for mouse clicks and eye tracking data that where further studied with respect to specific areas of interest (AOI). </p> <p>The dataset contains two screenshots indicating the areas of interest (AOIs; entitled “AreasOfInterest_Kartenportal.jpg and AreasOfInterest_Webportal.jpg), which separate the homepages into analyzable parts. It also contains eight screenshots of the homepages containing the total amount of collected clicks (each name starting with “clicks”) and two screenshots with the eye tracking heat maps (starting with “Heatmap”). The screenshots have directly been extracted from the click and eye tracking tools and matched with the before mentioned AOIs in order to gain the total count of clicks and views as well as the view duration the concerned area.</p> <p>All data are synthesized in a document containing three sheets with different tables: a first one with the initial data compilation for all AOIs of the two analyzed homepages (entitled “Data”), a second one with a more visual compiled data analysis for both homepages and all AOIs (entitled “Data2) and last one with the duration of the view as well as the duration of the fixation and the compiled click data (entitled « Eye tracking study data »).</p>
Mnemosine Digital Library Architecture
<p>Mnemosyne digital library architecture, identifying relationships with external systems.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.