Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
National Park Service - South Florida/Caribbean Inventory & Monitoring Network - SARI SET Surface Water level data from Salt River Bay National Historical Park and Ecological Preserve, St. Croix, US Virgin Islands.
Surface water level data (m) was collected in Salt River Bay National Historic Park and Ecological Preserve (SARI) by the South Florida/Caribbean Inventory and Monitoring Network (SFCN) as part of the Soil Elevation Table (SET) vital sign monitoring program. Water level data collected from 2017 to 2024 is included in this dataset. The water level data was collected using HOBOware Onset Water Level Data Loggers. This data-package is complete.
National Park Service - South Florida/Caribbean Inventory & Monitoring Network - Mary's Point SET Surface Water level data from Virgin Islands National Park, St. John, US Virgin Islands
Surface water level data (m) was collected in Virgin Islands National Park, Mary's Point (MARY) by the South Florida/Caribbean Inventory and Monitoring Network (SFCN) as part of the Soil Elevation Table (SET) vital sign monitoring program. Water level data collected from 2017 to 2024 is included in this dataset. The water level data was collected using HOBOware Onset Water Level Data Loggers. This data-package is complete.
National Park Service - South Florida/Caribbean Inventory & Monitoring Network - Water Creek SET Surface Water level data from Virgin Islands National Park, St. John, US Virgin Islands
Surface water level data (m) was collected in Virgin Islands National Park, Water Creek (WACR) by the South Florida/Caribbean Inventory and Monitoring Network (SFCN) as part of the Soil Elevation Table (SET) vital sign monitoring program. Water level data collected from 2017 to 2024 is included in this dataset. The water level data was collected using HOBOware Onset Water Level Data Loggers. This data-package is complete.
Survey of pond habitats and aquatic diversity of Madison, Wisconsin from May to Aug 2019 and 2020 Data Set
1) Urbanization may lead to changes in local richness (alpha diversity) or in community composition (beta diversity), although the direction of change can be challenging to predict. For instance, introduced species may offset the loss of native specialist taxa, leading to no change in alpha diversity in urban areas, but decreased beta diversity (i.e., more homogenous community structure). Alternatively, because urban areas can have low connectivity and high environmental heterogeneity between sites, they may support distinct communities from one another over small geographic distances. 2) Wetlands and ponds provide critical ecosystem services and support diverse communities, making them important systems in which to understand consequences of urbanization. To determine how urban development shapes pond community structure, we surveyed 68 ponds around Madison, Wisconsin, USA, which were classified as urban, greenspace, or rural based on surrounding land use. We evaluated the influence of local abiotic factors, presence of nonnative fishes, and landscape characteristics on alpha diversity of aquatic plants, macroinvertebrates, and vertebrates. We also analyzed whether surrounding land cover was associated with changes in community composition and/or the presence of specific taxa. 3) We found a 23% decrease in mean richness (alpha diversity) from rural to urban pond sites, and a 15% decrease in richness from rural to urban greenspace pond sites. Among landscape factors, observed pond richness was negatively correlated with adjacent developed land and mowed lawns, as well as greater distances to other waterbodies. Among pond level factors, habitat complexity was associated with increased richness, while the presence of invasive fish was associated with decreased richness. 4) Beta diversity was relatively high for all ponds due to turnover in composition between sites. Urban ponds supported more introduced species, lacked a subset of native species found in rural ponds,
Surface Elevation Table (SET) data (pin heights) from control and fertilized plots at Spartina alterniflora, S. patens and Typha sp. marshes, Plum Island Ecosystems LTER, MA (1999-2025).
A Surface Elevation Table (SET) is used to measure changes in the elevation of the marsh surface at three long term marsh fertilization experimental research sites. The sites include one Typha-dominated brackish marsh, one Spartina alterniflora-dominated salt marsh, and one S. patens-dominated salt marsh. Sites are located on the Rowley and upper Parker Rivers in the Plum Island Ecosystems (PIE) LTER site.
Surface Elevation (SET) Data for Brownsville Forest at the Virginia Coastal Reserve, 2019-2025
This dataset contains data from Surface Elevation Tables (SETs) located in the Brownsville Forest at the Virginia Coastal Reserve. These research plots were set up in 2019 as a part of a larger forest disturbance project. Marker Horizons and Shallow SETs were installed in 2020 but have not yet been measured. Tyler Messerschmidt has made all measurements of these SETs
Software Developer Expertise GitHub and Stack Overflow data sets
<p>Cross-Platform Software Developer Expertise Learning by Norbert Eke</p> <p>This data set is part of my Master's thesis project on developer expertise learning by mining Stack Overflow (SOTorrent) and Github (GHTorrent) data. Check out my portfolio website at norberte.github.io</p>
Data set for Global quantitative synthesis of ecosystem functioning across climatic zones and ecosystem types
<p>Dataset used in the publication: " Global quantitative synthesis of ecosystem functioning across climatic zones and ecosystem types". The dataset gathers estimates of ecosystem standing stocks (biomass, organic carbon, detritus), fluxes (GPP, ER, NEP) and process rates (decomposition and carbon uptake rates) for eight broad ecosystem types (forest, grassland, agroecosystem, desert, stream, lake, pelagic and benthic marine ecosystems) in five broad climatic zones (arctic, boreal, arid, temperate, tropical, arid).</p> <p>The scripts to produce the figures and the statistics of the publication are released along with the txt version of the data, which file is uploaded when running the script.</p>
Obtaining Problem Statements and Transformation into Ideas - Workshop Data Set
<p>In order to enable members of a socio-technical evolutionary-teal organization to design their technical component, we conducted a workshop that structures the collaboration between technical trained participants and non-trained participants. F general challenges in the form of “problem statements” have been gathered. Afterwards, these are transformed into ideas for organizational and technical change.</p> <p>The design thinking workshop is set up as a standalone, roughly five-hour group discussion, following the principles of participatory design. It has been recorded in video and this data set contains a textual German transcript and transcripts of the moderation cards that have been created during the workshop.</p> <p>We hope that the material can be used to (a) comprehend the interpretation used in our qualitative research, (b) to adapt the workshop model by other volunteers of our case study Viva con Agua de St. Pauli e.V. (https://www.vivaconagua.org/), and (c) investigate other interesting research questions.</p>
Thematic/Taxonomic Analogy Task Data Set
<p>A classic analogy paradigm (A:B::C:?) was developed to test children's understanding of thematic and taxonomic categorization of offd in children aged between 3 and 6 years old. Participants' level of food rejection disposition was also measured using the Child Food Rejection Scale (CFRS; Rioux, Lafraire, Picard, 2017) to determine how food rejection affects children's categorization ability in the food domain. </p> <p>Data sets for: children's food rejection scores, children's responses for thematic and taxonomic conditions of food categorization analogy task, and children's responses for stimuli identification task</p> <p>Supplemental material: Stimuli overview</p>
Labelled magnetic reconnection simulation data set
<p>Numerical simulations have been performed on Marconi at CINECA (Italy) under the ISCRA initiative. The corresponding data can be found at: <a href="https://doi.org/10.5281/zenodo.3935887">https://doi.org/10.5281/zenodo.3935887</a></p>
Data set and code supporting Marshall et al. 2020. No room to roam: King Cobras reduce movement in agriculture.
<p>Data and code used in the publication:</p> <p>Marshall, B.M., Crane, M., Silva, I., Strine, C.T., Jones, M.D., Hodges, C.W., Suwanwaree, P., Artchawakom, T., Waengsothorn, S., Goode, M. (2020). No room to roam: King Cobras reduce movement in agriculture. <em>Mov Ecol</em> <strong>8, </strong>33 (2020). https://doi.org/10.1186/s40462-020-00219-5</p> <p>Marshall, B.M., Crane, M., Silva, I., Strine, C.T., Jones, M.D., Hodges, C.W., Suwanwaree, P., Artchawakom, T., Waengsothorn, S., Goode, M. (2020). No room to roam: King Cobras reduce movement in agriculture. bioRxiv 2020.03.24.006676; doi: https://doi.org/10.1101/2020.03.24.006676</p> <p>Including: telemetry data, habitat shapefile and derived rasters, ISSF and JAGS model specification and results, code to reproduce analysis and generate figures. </p>
The NANOGrav 12.5-year Narrowband Data Set (version 12yv4)
<p>The NANOGrav 12.5-year narrowband data set (public release "12yv4") is the supplemental data set accompanying Alam et al. 2021, "The NANOGrav 12.5 yr Data Set: Observations and Narrowband Timing of 47 Millisecond Pulsars," 2021, Astrophysical Journal Supplements, 252, 4, DOI: 10.3847/1538-4365/abc6a0. It contains narrowband pulse times of arrival, one-dimensional pulse templates, pulsar timing models, timing residuals, and clock files.</p> <p>Details about the contents of these files are contained in NANOGrav_12yv4_narrowband/README, as well as in NANOGrav_12yv4_narrowband/narrowband/README.narrowband. The wideband version of this dataset (published in Alam et al. 2021, ApJS, 252, 5, DOI: 10.3847/1538-4365/abc6a1) can be found at Zenodo DOI: 10.5281/zenodo.4312887. Both the narrowband and wideband data sets are also available at data.nanograv.org.</p> <p>Analysis of the 12.5-year narrowband dataset will be published in Arzoumanian et al. 2021, "The NANOGrav 12.5-year Data Set: Search for an Isotropic Gravitational-Wave Background," accepted for publication in ApJ Letters, 2020arXiv200904496A.</p>
Data set and code supporting Marshall et al., "An inventory of online reptile images"
<p>Data set and code supporting: MARSHALL, B.M., FREED, P., VITT, L.J., BERNARDO, P., VOGEL, G., LOTZKAT, S., FRANZEN, M., HALLERMANN, J., SAGE, R.D., BUSH, B. and DUARTE, M.R., 2020. An inventory of online reptile images. <em>Zootaxa</em>, <em>4896</em>(2), pp.251-264. DOI:<a href="https://doi.org/10.11646/zootaxa.4896.2.6">10.11646/zootaxa.4896.2.6</a></p> <p>Data includes: </p> <ul> <li>Supplementary Table 1. List of all species and the number of photos in each of the 6 repositories: "SuppData1_Species_Photo_Count_Table_2020-08-04_no_syn.csv"</li> <li>Supplementary Table 2. List of species without photo in any of the 6 repositories: "SuppData2_Species_no_photos.csv"</li> <li>Supplementary Table 3. Per country summary data of number of species present and number with images: "SuppData3_Country_species_counts.csv"</li> <li>Reptile Database species checklist: "reptile_checklist_2020_04.csv"</li> <li>Reptile Database species synonyms used in second Wikimedia search: "reptile names 2019 syno.csv"</li> </ul> <p>Code includes:</p> <ul> <li>R code used to retrieve Flickr photograph metadata: "SuppCode1_Flickr_search.R"</li> <li>R code used to retrieve Wikimedia photograph metadata: "SuppCode2_Wikimedia_query.R"</li> <li>R code used to retrieve HerpMapper photograph metadata: "SuppCode3_HerpMapper_search.R"</li> <li>R code used to generate figures: "SuppCode4_Figure Generation.R"</li> </ul> <p>Also includes Zootaxa supplementary table.</p> <p> </p>
NAPv1.0: A seasonal hydrographic gridded data set for the Northern Antarctic Peninsula, Southern Ocean
<p>The Northern Antarctic Peninsula (NAP) climatology version 1 (NAPv1.0) was built by optimally interpolate hydrographic data sets from the CTD, MEOP and Argo floats profiles sampled in the NAP and adjacent regions during the period of 1990-2019. The database consists of data from the World Ocean Database, Pangaea, Hutchinson et al. (2020), Brazilian High Latitude Oceanography Group (GOAL; http://goal.furg.br/), Marine Mammals Exploring the Oceans Pole to Pole consortium (MEOP), and Argo floats. The climatology has outputs for summer (Jan-Mar), autumn (Apr-Jun), winter (Jul-Sep) and spring (Oct-Dec). The profiles were first linearly interpolated onto 90 depth levels, and then optimally interpolated in space using a grid of ~10 km resolution. The grid spacing is 0.09˚ along latitudes and 0.2˚ along longitudes (i.e., 0.09˚ latitude x 0.09˚/cos(63˚S) longitude, where 63˚S is the mean latitude of our domain). A series of tests were made to find the appropriate smoothing lengthscale and the a priori relative error in order to find a balance between smoothness and feature representativeness. The final smoothing lengthscale (i.e. the radius of influence of the interpolation) chosen was 1˚ in latitude and longitude, and the a priori relative error allowed was set to 0.2 for the objective interpolation algorithm. The same constants were set for all depth levels and all variables. The regions where the mapping relative error was higher than 0.5 were excluded. The NAPv1.0 climatology can be used for several applications, including input data for ocean and climate models initialization/assessment and ocean reanalysis evaluation, as well as to produce and reconstruct biogeochemical properties. The NAPv1.0 climatology represents the ocean mean-state for the NAP for the end of the 20th and early 21st-century.</p> <p> </p> <p><strong>Reference: </strong><br> Dotto, T. S., Mata, M. M., Kerr, R., and Garcia, C. A. E.: A novel hydrographic gridded data set for the northern Antarctic Peninsula, Earth Syst. Sci. Data, 13, 671–696, https://doi.org/10.5194/essd-13-671-2021, 2021.</p>
Phindr3D: Test Data Set 1 (primary mouse cortical neurons)
<p>3D confocal image stacks of primary cortical neurons under different treatment conditions to test the functionality of Phindr3D. Explanatory .txt file contained in the ZIP file.</p> <p>Please see the manuscript for details and on how to access the full data set:</p> <p> </p> <p><strong>Rapid 3D phenotypic analysis of neurons and organoids using data-driven cell segmentation-free machine learning</strong></p> <p>Philipp Mergenthaler*, Santosh Hariharan*, James M. Pemberton, Corey Lourenco, Linda Z. Penn, David W. Andrews</p> <p><em>PLOS Computational Biology, DOI: <a href="https://dx.doi.org/10.1371/journal.pcbi.1008630">10.1371/journal.pcbi.1008630</a></em></p> <p> </p> <p><strong>Phindr3D is available on GitHub</strong>: <a href="https://github.com/DWALab/Phindr3D">GitHub - DWALab/Phindr3D</a></p> <p> </p>
Phindr3D: Test Data Set 2 (human MCF10A breast cancer organoids)
<p>3D confocal image stacks of human MCF10A breast cancer organoids expressing different oncogenes to test the functionality of Phindr3D. Explanatory .txt file contained in the ZIP files.</p> <p>Please see the manuscript for details and on how to access the full data set:</p> <p> </p> <p><strong>Rapid 3D phenotypic analysis of neurons and organoids using data-driven cell segmentation-free machine learning</strong></p> <p>Philipp Mergenthaler*, Santosh Hariharan*, James M. Pemberton, Corey Lourenco, Linda Z. Penn, David W. Andrews</p> <p><em>PLOS Computational Biology, DOI: <a href="https://dx.doi.org/10.1371/journal.pcbi.1008630">10.1371/journal.pcbi.1008630</a></em></p> <p> </p> <p><strong>Phindr3D is available on GitHub</strong>: <a href="https://github.com/DWALab/Phindr3D">GitHub - DWALab/Phindr3D</a></p>
Data set supplementing "Benchmarking triage capability of symptom checkers against that of medical laypersons: Survey study"
<p>This is the de-identified data set used to conduct the analyses in the study published as Original Research in the JMIR under the title "Benchmarking triage capability of symptom checkers against that of medical laypersons: Survey study" (https://doi.org/10.2196/24475)</p> <p>The data set contains the assessments of the urgency of symptoms to 45 fictitious clinical case vignettes by 91 US participants, and the participants' age, gender and level of education. Data for the symptom checker apps is needed to fully reproduce our study and can be found in the appendix of the paper "Evaluation of symptom checkers for self diagnosis and triage: audit study" by Semigran et al. (2015) (https://doi.org/10.1136/bmj.h3480).</p>
TCOM-HF : Daily global gap-free stratospheric hydrogen fluoride (HF) profile data set based on TOMCAT CTM and Occultation Measurements
<p><strong>Methodology: TOMCAT simulation is performed at T64L32 resolution for the 2000-2024 time period. Collocated hydrogen fluoride (HF) profiles are divided in five latitude bins: SH polar (90S-50S), SH mid-lat (70S-20S), tropics (40S-40N), NH mid-lat (20N-70N) and NH polar (50N-90N). Initially, model-measurement differences are calculated for each zonal bins (51 height levels, 10km to 60km). Note that if enough ACE measurements are not avaliable for a particular level then data is purely based on TOMCAT simulated output field. Separate XGBoost regression models are trained for the differences between TOMCAT and measurements at each level for a given latitude bin. XGBoost model is then used to estimate error corrections for all the TOMCAT grids. TOMCAT output sampled at 1.30 pm local time at the equator. Estimated corrections for a given model grid that are added to the original TOMCAT simulated day and night time hydrogen fluoride profiles. Height resolved data are then interpolated on 28-pressure levels (300 - 0.1hPa). For overlapping latitude bins, we use averages and then calculate daily zonal mean values. For more details see attached presentation. Previous version use both HALOE and ACE data. Here only ACE data is used.</strong></p> <p><strong>Dataset also includes two files containing daily mean zonal mean hydrogen fluoride profiles on height (10-50 km) and pressure (300-0.1 hPa) levels:</strong></p> <p><strong>zmhf_TCOM_hlev_T2Dz_2000_2024.nc – height level data (10 to 50 km)</strong></p> <p><strong>zmhf_TCOM_plev_T2Dz_2000_2024.nc – pressure level data (300 to 0.1 hPa)</strong></p> <p><strong>Daily 3D profiles on height and pressure levels would be made available on request.</strong></p>
MOSAiC Cloudnet issue data set
<p>This data set contains information on possible data issues caused by external drivers (e.g. tethered balloon artefacts in the observations) related to the MOSAiC Cloudnet data set. Flagged data must be handled with care and should be excluded from statistical analyses. Issues tracking flags are identified by tethered balloon operation periods and experienced-eye observations of MOSAiC staff.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.