Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

95

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

95 results for “Digital collections”

Learn how ShareScore rates datasets ↗
zenodo40/100

Berlin State Library (2023). Fulltexts of the Digitized Collections of the Berlin State Library (SBB)

<p>The motivation for creating this dataset was to enable research on the basis of fulltexts which are available in a cultural heritage institution on a large scale. Libraries such as the Berlin State Library (SBB) typically provide three kinds of data: Images (scans of books, illustrations contained in the scanned material, or else), metadata (descriptive data providing information on the digitized item) and texts. The latter are usually received via an implementation of optical character recognition (OCR) of the digitized books or manuscripts. In the <a href="https://digital.staatsbibliothek-berlin.de/">digitized collections of the Staatsbibliothek zu Berlin (SBB)</a>, the fulltexts can be downloaded manually, item by item. The publication of a set of about 5 million OCR&rsquo;d pages alleviates the accessibility of the fulltexts. The basic funding interest in this data publication is the stimulation of innovation.</p> <p>The dataset consists of a single sqlite database containing ALL the fulltexts available in the digitized collections of the Berlin State Library as of August 21st, 2019, with information on languages and entropy added on March 1st, 2023. The size of the sqlite file is about 15.8 GB.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Digitized geological and geophysical data from the Po Plain and the northern Adriatic Sea (north Italy) collected from public sources.

<p>The database is a supplementary material of:</p><p>Livani, M., Petracchini, L., Benetatos, C., Marzano, F., Billi, A., Carminati, E., Doglioni, C., Petricca, P., Maffucci, R., Codegone, G., Rocca, V., Antoncecchi, I. (2023). Subsurface geological and geophysical data from the Po Plain and the northern Adriatic Sea (north Italy). Earth System Science Data Discussions, 15, n. 9,&nbsp; 4261-4293, doi: 10.5194/essd-15-4261-2023</p><p>&nbsp;</p><p>This database comprises subsurface geological and geophysical data from the Po Plain and the northern Adriatic Sea (north Italy). We realized the database by collecting, revising, and digitizing data, originally in raster format, from public sources. These data have been then used to reconstruct the overall subsurface 3D architecture and to extract the physical properties of the subsurface geological units.</p><p>The data have a common geographical system: WGS 84/UTM zone 32N; EPSG: 32632.</p><p>The database contains borehole information from 160 deep wells (i.e., wellhead coordinates, rotary table elevation, measured depth, true depth, total depth and deviation survey) and digitized Spontaneous Potential, Gamma Ray, and Sonic logs. Five horizons were digitized from 61 geological cross-sections and from 10 isobath maps that roughly correspond to the boundaries of units showing different lithological properties and with different mechanical properties. The horizons are, from the oldest to the youngest: the top of the magnetic basement, the top of the carbonate succession, the base of the Pliocene, the base of the Calabrian and the base of recent continental deposits. In addition, the gridded surfaces of the 3D geological model are available.</p><p>We organized the database into two groups: "primitive data" and "derived data".</p><p>Primitive data are the result of the digitization of public data. The database of the primitive data is formed by the main horizons reported in the geological cross-sections, the isobaths of the main geological surfaces, the well locations, comprised their trajectory along depth, lithological and stratigraphical information, and geophysical logs from composite well logs. Well data include specific sets of well logs aimed to geological/mechanical characterization of the geological units (e.g., Spontaneous Potential log, Resistivity log, Gamma Ray log, and Sonic log).</p><p>Derived data consist of two datasets: i) the primitive isobaths maps and geological cross-sections data that have been filtered and verified after a data accuracy analysis performed to unravel discrepancies in the interpretation of the subsurface geological horizons; ii) a set of regional surfaces of the main geological units of the Po Plain subsurface. These surfaces were generated by interpolating the filtered primitive data and without considering the fault occurrence/displacements.</p><p>Primitive and derived data are provided in delimited text file format organized according to the data type (i.e., well, geological cross-section, map and gridded surface). The format and the organization of data are explained in the related "readme" file presents within each data folder.</p><p>Our database represents a collection of the main published works regarding the Po Plain. Detailed studies related to specific sectors of the Po Plain might not be present in our database.</p><p>Further details about the data processing and organization are given in the related manuscript.</p><p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

"The sound comes from a meadow in the Sierra Nevada Mountains in California. The meadow is at an elevation of 2400 meters near a mountain named Olancha Peak, which is 3700 meters in altitude. Ihave a group of friends with which Ibackpack (trek) into the mountains. Our goal was to spend some time in the mountains and hike to the top of Olancha Peak (…) By the time we reached the meadow, we were in a forest and there was still snow on the ground in some places. We took the trip in June of 2006. The Sierra Nevada Mountains are a large mountain range. Much of the range is protected by national parks or preserved areas we call 'wilderness areas' (…) Ihave been backpacking for nearly 40 years and Iwill hopefully continue with this challenging activity for 40 years more! Many of my friends are much younger than Iam and it gives me much satisfaction to be able to have as much or more stamina for this activity than they have! When we are on these trips, we hike up peaks, catch fish, drink some whiskey around campfires and enjoy our time in the beautiful solitude. My memories of this trip were of the steep, hot hike from the desert to the cool meadow; the overall beauty of the nature, the absolute solitude of our campsite near the meadow; the strenuous hike to the top of Olancha Peak; the camaraderie of my friends; and, of course the sound of the frogs in the meadow. The frog sounds were astounding to me and Iwould listen in awe of the creature's instinctual desire to reproduce and continue the existence of their kind. Surely there were different species in the meadow for some of the frog sounds were different than others. The sounds only occurred after the Sun went down for the evening. Istood next to the creek in the meadow and recorded the sounds using my digital camera." [Peter/plentz1960]16 in Collecting Sounds. Online Sharing of Field Recordings as Cultural Practice

"The sound comes from a meadow in the Sierra Nevada Mountains in California. The meadow is at an elevation of 2400 meters near a mountain named Olancha Peak, which is 3700 meters in altitude. Ihave a group of friends with which Ibackpack (trek) into the mountains. Our goal was to spend some time in the mountains and hike to the top of Olancha Peak (…) By the time we reached the meadow, we were in a forest and there was still snow on the ground in some places. We took the trip in June of 2006. The Sierra Nevada Mountains are a large mountain range. Much of the range is protected by national parks or preserved areas we call 'wilderness areas' (…) Ihave been backpacking for nearly 40 years and Iwill hopefully continue with this challenging activity for 40 years more! Many of my friends are much younger than Iam and it gives me much satisfaction to be able to have as much or more stamina for this activity than they have! When we are on these trips, we hike up peaks, catch fish, drink some whiskey around campfires and enjoy our time in the beautiful solitude. My memories of this trip were of the steep, hot hike from the desert to the cool meadow; the overall beauty of the nature, the absolute solitude of our campsite near the meadow; the strenuous hike to the top of Olancha Peak; the camaraderie of my friends; and, of course the sound of the frogs in the meadow. The frog sounds were astounding to me and Iwould listen in awe of the creature's instinctual desire to reproduce and continue the existence of their kind. Surely there were different species in the meadow for some of the frog sounds were different than others. The sounds only occurred after the Sun went down for the evening. Istood next to the creek in the meadow and recorded the sounds using my digital camera." [Peter/plentz1960]16

opencc-by-4.0Dec 2019View details →
dryad40/100

Data from: Integrating deep learning derived morphological traits and molecular data for total-evidence phylogenetics: lessons from digitized collections

Open the record for dataset details and reuse information.

publicDec 2024View details →
zenodo36/100

Going digital: Added value of electronic data collection in 2018 Afghanistan Health Survey

<p>These data were collected and analyzed to compare measures of cost efficiency, data quality and user acceptability between paper-based and digitally collected data for a nationally representative household survey in Afghanistan.&nbsp;</p>

opencc-by-4.0Nov 2020View details →
dryad36/100

Data from: Quantifying the dark data in museum fossil collections as palaeontology undergoes a second digital revolution

Large-scale analysis of the fossil record requires aggregation of palaeontological data from individual fossil localities. Prior to computers these synoptic datasets were compiled by hand, a laborious undertaking that took years of effort and forced palaeontologists to make difficult choices about what types of data to tabulate. The advent of desktop computers ushered in palaeontology's first digital revolution – online literature-based databases, such as the Paleobiology Database (PBDB). However, the published literature represents only a small proportion of the palaeontological data housed in museum collections. Although this issue has long been appreciated, the magnitude, and thus potential significance, of these so-called "dark data" has been difficult to determine. Here, in the early phases of a second digital revolution in palaeontology the digitization of museum collections – we provide an estimate of the magnitude of palaeontology's dark data. Digitization of our nine institutions' holdings of Cenozoic marine invertebrate collections from California, Oregon, and Washington in the United States reveals that they represent 23 times the number of unique localities than are currently available in the Paleobiology Database. These data, and the vast quantity of similarly untapped dark data in other museum collections, will when digitally mobilized enhance palaeontologists' ability to make inferences about the patterns and processes of past evolutionary and ecological changes.

opencc-zeroDec 2017View details →
dryad36/100

Data from: A new digital method of data collection for spatial point pattern analysis in grassland communities

<p>A major objective of plant ecology research is to determine the underlying processes responsible for the observed spatial distribution patterns of plant species. Plants can be approximated as points in space for this purpose, and thus, spatial point pattern analysis has become increasingly popular in ecological research. The basic piece of data for point pattern analysis is a point location of an ecological object in some study region. Therefore, point pattern analysis can only be performed if data can be collected. However, due to the lack of a convenient sampling method, a few previous studies have used point pattern analysis to examine the spatial patterns of grassland species. This is unfortunate because being able to explore point patterns in grassland systems has widespread implications for population dynamics, community-level patterns and ecological processes. In this study, we develop a new method to measure individual coordinates of species in grassland communities. This method records plant growing positions via digital picture samples that have been sub-blocked within a geographical information system (GIS). Here, we tested out the new method by measuring the individual coordinates of <i>Stipa</i><i> grandis</i> in grazed and ungrazed <i>S. grandis</i> communities in a temperate steppe ecosystem in China. Furthermore, we analyzed the pattern of <i>S. grandis</i> by using the pair correlation function <i>g</i>(<i>r</i>) with both a homogeneous Poisson process and a heterogeneous Poisson process. Our results showed that individuals of <i>S. grandis</i> were overdispersed according to the homogeneous Poisson process at 0-0.16 m in the ungrazed community, while they were clustered at 0.19 m according to the homogeneous and heterogeneous Poisson processes in the grazed community. These results suggest that competitive interactions dominated the ungrazed community, while facilitative interactions dominated the grazed community. In sum, we successfully executed a new sampling method, using digital photography and a Geographical Information System, to collect experimental data on the spatial point patterns for the populations in this grassland community.</p>

opencc-zeroJun 2021View details →
zenodo36/100

Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center

<p>This is the Tibetan etext collection of the Buddhist Digital Resource Center (www.tbrc.org) as of April 28, 2017.</p>

opencc-by-4.0Apr 2017View details →
zenodo36/100

Digital Library Mnemosine. Workflow for creating new collections

<p><span>Flow chart for semi-automatic creation of new collections in Mnemosine.</span></p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

FIGURE 2 in Digitization workflows for paleontology collections

FIGURE 2. Example implementation of the workflow modules described herein.

opencc-by-4.0Oct 2016View details →
dryad36/100

Data from: A new digital method of data collection for spatial point pattern analysis in grassland communities

Open the record for dataset details and reuse information.

publicJul 2021View details →
dryad36/100

Data from: Quantifying the dark data in museum fossil collections as palaeontology undergoes a second digital revolution

Open the record for dataset details and reuse information.

publicAug 2018View details →
zenodo32/100

Addressing bias in digital cultural heritage collections metadata: the example of the DE-BIAS project

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Access to Collections and Digital Preservation

<p>Audio from&nbsp;<em>Session 9: Access to Collections and Digital Preservation</em>, held Thursday 27 June&nbsp;2019 at the LIBER 2019 Annual Conference.</p> <p>Talks included:</p> <ul> <li> <p>9.1&nbsp;<a href="https://doi.org/10.5281/zenodo.3259715">Access to Collections: an Essential Part of Research Collaborations</a>, Alex Fenlon, University of Birmingham, United Kingdom</p> </li> <li> <p>9.2&nbsp;<a href="https://doi.org/10.5281/zenodo.3259719">Clear and Consistent: Copyright Assessment Framework for Libraries</a>, Fred Saunderson, National Library of Scotland, United Kingdom, Dafydd Tudur, National Library of Wales, United Kingdom</p> </li> <li> <p>9.3&nbsp;<a href="https://doi.org/10.5281/zenodo.3259721">Networking with Networks: What is the Landscape for Digital Preservation Communities like?</a>&nbsp;Thomas B&auml;hr and Michelle Lindlar, TIB Leibniz Information Center for Science and Technology University Library, Germany, Sabine Schrimpf, Deutsche Nationalbibliothek, Germany, Stefan Strathmann, Staats- und Universit&auml;tsbibliothek G&ouml;ttingen, Germany, Monika Zarnitz, ZBW Leibniz-Information Center for Economics, Germany</p> </li> </ul>

opencc-by-4.0Jul 2019View details →
zenodo32/100

Barbara Thiers of the New York Botanical Garden and president of the Society for the Preservation of Natural History Collections is among those leading the effort to harness the explosion of data being digitized by collections around the globe. Photograph: New York Botanical Garden, Bronx, NY. in The Evolution of Natural History Collections

Barbara Thiers of the New York Botanical Garden and president of the Society for the Preservation of Natural History Collections is among those leading the effort to harness the explosion of data being digitized by collections around the globe. Photograph: New York Botanical Garden, Bronx, NY.

opennotspecifiedMar 2019View details →
ClinicalTrials.gov32/100

Study to Learn More About Outcomes Reported by Patients Suffering From Type 2 Diabetes Using a Digital Data Collection Tool

ClinicalTrials.gov study NCT04383041. IPD Sharing: UNDECIDED. Countries: 1. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

A Remote, 9-week Insomnia Treatment Trial to Collect Real World Data for a Digital Therapeutic

ClinicalTrials.gov study NCT04325464. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Unhide® Project: A Digital Health Platform to Collect Lifestyle Data for Brain Inflammation Research

ClinicalTrials.gov study NCT04806620. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
dryad32/100

Data from: Using digitized museum collections to understand the effects of habitat on wing coloration in the Puerto Rican monarch

Open the record for dataset details and reuse information.

publicJun 2019View details →
dryad28/100

Data from: A survey of digitized data from U.S. fish collections in the iDigBio data aggregator

Recent changes in institutional cyberinfrastructure and collections data storage methods have dramatically improved accessibility of specimen-based data through the use of digital databases and data aggregators. This analysis of digitized fish collections in the U.S. demonstrates how information from data aggregators, in this case iDigBio, can be extracted and analyzed. Data from U.S. institutional fish collections in iDigBio were explored through a strictly programmatic approach using the ridigbio package and fishfindR web application. iDigBio facilitates the aggregation of collections data on a purely voluntary fashion that requires collection staff to consent to sharing of their data. Not all collections are sharing their data with iDigBio, but the data harvested from 38 of the 143 known fish collections in the U.S. that are in iDigBio account for the majority of fish specimens housed in U.S. collections. In the 22 years since publication of the last survey providing information on these 38 collections, 1,219,168 specimen records (lots), 15,225,744 specimens, 3,192 primary types, and 32,868 records of secondary types have been added. This is an increase of 65.1% in the number of cataloged records and an increase of 56.1% in the number of specimens. In addition to providing specimen-based data for research, education, and various outreach activities, data that are accessible via data aggregators can be used to develop accurate, up-to-date reports of information on institutional collections. Such reports present collections data in an organized and accessible fashion and can guide targeted efforts by collections personnel to meet discipline-specific needs and make data more transparent to downstream users. Data from this survey will be updated and published regularly in a dynamic web application that will aid collections staff in communicating collections value while simultaneously giving stakeholders a way to explore collections holdings as they relate to the institutions in which they are housed. It is through this resource that collections will be able to leverage their data against those of similar collections to aid in the procurement of financial and institutional support.

opencc-zeroDec 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record