Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8,819
datasets available to search
ShareScore release 0.7.1
Dataset results
8,819 results for “new data”
Monthly precipitation data from a network of standard gauges at the Jornada Experimental Range (Jornada Basin LTER) in southern New Mexico, January 1916 - ongoing
This ongoing dataset contains monthly precipitation measurements from a network of standard can rain gauges at the Jornada Experimental Range in Dona Ana County, New Mexico, USA. Precipitation physically collects within gauges during the month and is manually measured with a graduated cylinder at the end of each month. This network is maintained by USDA Agricultural Research Service personnel. This dataset includes 39 different locations but only 29 of them are current. Other precipitation data exist for this area, including event-based tipping bucket data with timestamps, but do not go as far back in time as this dataset.
Mammal occurrence data derived from camera traps in grassland-shrubland ecotones at 24 sites in the Jornada Basin, southern New Mexico, USA, 2014-ongoing
The objective of this ongoing study is to investigate how abundance, distribution, and activity of mammals (>= 1 kg) vary across grassland to shrubland ecotones in the northern Chihuahuan Desert. This dataset includes animal occurrence data derived from camera trap images captured in 24 grassland-to-shrubland ecotone sites in the Jornada Basin, Dona Ana County, New Mexico, USA. The data set contains occurrence records from 14 mammal species with the date and time a species was detected. Also included are the number of individuals in a photo, operational dates and number of functional camera days for cameras, total number of trap nights a camera was active, and geographical coordinates of camera trap locations. Sampling is ongoing and occurs during the monsoon season from July-November. Sampling has occurred annually since 2014.
Meteorology Data from the Sevilleta National Wildlife Refuge, New Mexico
These files contain hourly meteorological data that were collected from a network of permanent weather stations on the Sevilleta National Wildlife Refuge as part of the Sevilleta Long Term Ecological Research Program.
Net Primary Productivity (NPP) Weight Data at the Sevilleta National Wildlife Refuge, New Mexico
Several long-term studies at the Sevilleta LTER measure net primary production (NPP) across ecosystems and treatments. Net primary production is a fundamental ecological variable that quantifies rates of carbon consumption and fixation. Estimates of NPP are important in understanding energy flow at a community level as well as spatial and temporal responses to a range of ecological processes. Above-ground net primary production (ANPP) is the change in plant biomass, including loss to death and decomposition, over a given period of time. To measure this change, vegetation variables, including species composition and the cover and height of individuals, are sampled up to three times yearly (winter, spring, and fall) at permanent plots within a study site. The weight data presented here is obtained by harvesting a series of covers for species observed during plot sampling. These species are always harvested from habitat comparable to the plots in which they were recorded. This data is then used to make volumetric measurements of species and build regressions correlating biomass and volume. From these calculations, seasonal biomass and seasonal and annual NPP are determined.
Gunnison's Prairie Dog Restoration Experiment (GPDREx): Vegetation Cover Data from the Sevilleta National Wildlife Refuge, New Mexico (2011-2016)
Prairie dogs (Cynomys spp.) are burrowing rodents considered to be ecosystem engineers and keystone species of the central grasslands of North America. Yet, prairie dog populations have declined by an estimated 98% throughout their historic range. This dramatic decline has resulted in the widespread loss of their important ecological role throughout this grassland system. The 92,060 ha Sevilleta NWR in central New Mexico includes more than 54,000 ha of native grassland. Gunnison's prairie dogs (C. gunnisoni) were reported to occupy ~15,000 ha of what is now the SNWR during the 1960's, prior to their systematic eradication. In 2010, we collaborated with local agencies and conservation organizations to restore the functional role of prairie dogs to the grassland system. Gunnison's prairie dogs were reintroduced to a site that was occupied by prairie dogs 40 years ago. This work is part of a larger, long-term study where we are studying the ecological effects of prairie dogs as they re-colonize the grassland ecosystem.
Core Site Grid Quadrat Data for the Net Primary Production Study at the Sevilleta National Wildlife Refuge, New Mexico
Begun in spring 2013, this project is part of a long-term study at the Sevilleta LTER measuring net primary production (NPP) across three distinct ecosystems: creosote-dominant shrubland (Site C), black grama-dominant grassland (Site G), and blue grama-dominant grassland (Site B). Net primary production is a fundamental ecological variable that quantifies rates of carbon consumption and fixation. Estimates of NPP are important in understanding energy flow at a community level as well as spatial and temporal responses to a range of ecological processes. Above-ground net primary production is the change in plant biomass, represented by stems, flowers, fruit and foliage, over time and incorporates growth as well as loss to death and decomposition. To measure this change the vegetation variables in this dataset, including species composition and the cover and height of individuals, are sampled twice yearly (spring and fall) at permanent 1m x 1m plots within each site. A third sampling at Site C is performed in the winter. The data from these plots is used to build regressions correlating biomass and volume via weights of select harvested species obtained in SEV999, "Net Primary Productivity (NPP) Weight Data." This biomass data is included in SEV999, "Seasonal Biomass and Seasonal and Annual NPP for Core Grid Research Sites."
Monsoon Rainfall Manipulation Experiment (MRME) Soil Temperature, Moisture and Carbon Dioxide Data from the Sevilleta National Wildlife Refuge, New Mexico
The Monsoon Rainfall Manipulation Experiment (MRME) is designed to understand changes in ecosystem structure and function of a semiarid grassland caused by increased precipitation variability, by altering rainfall pulses, and thus soil moisture, that drive primary productivity, community composition, and ecosystem functioning. The overarching hypothesis being tested is that changes in event size and frequency will alter grassland productivity, ecosystem processes, and plant community dynamics. Treatments include (1) a monthly addition of 20 mm of rain in addition to ambient, and a weekly addition of 5 mm of rain in addition to ambient during the months of July, August and September. It is predicted that changes in event size and variability will alter grassland productivity, ecosystem processes, and plant community dynamics. In particular, we predict that many small events will increase soil CO2 effluxes by stimulating microbial processes but not plant growth, whereas a small number of large events will increase aboveground NPP and soil respiration by providing sufficient deep soil moisture to sustain plant growth for longer periods of time during the summer monsoon.
SEV-LTER Mean - Variance Experiment Seasonal Biomass Data at the Sevilleta National Wildlife Refuge, New Mexico
We designed novel field experimental infrastructure to resolve the relative importance of changes in the climate mean and variance in regulating the structure and function of dryland populations, communities, and ecosystem processes. The Mean - Variance Climate Experiment (MVE) adds three novel elements to prior designs that have manipulated interannual variance in climate in the field (Gherardi & Sala, 2013) by (i) determining interactive effects of mean and variance with a factorial design that crosses reduced mean with increased variance, (ii) studying multiple dryland biomes to compare their susceptibility to transition under interactive climate drivers, and (iii) adding stochasticity to our treatments to permit the antecedent effects that occur under natural climate variability. This new infrastructure enables direct experimental tests of the hypothesis that interactions between the mean and variance of precipitation will have larger ecological impacts than either the mean or variance in precipitation alone. This data package includes species-level plant cover and biomass data from the Mean - Variance Experiment at five sites comprising the major ecosystems of the Sevilleta National Wildlife Refuge: Chihuahuan Desert shrubland, Chihuahuan Desert grassland, Great Plains grassland, Juniper savanna, and pinon-juniper woodland. Species cover and volume in one-meter-squared quadrats are assessed twice-yearly in spring and fall, and regressions correlating biomass and volume constructed using seasonal harvest weights from SEV157, "Net Primary Productivity (NPP) Weight Data."
Data: Algorithms for new types of fair stable matchings
<p>This data corresponds to the data and experiments described in Section 5 of<br> the following paper:</p> <p>Algorithms for new types of fair stable matchings<br> Authors: Frances Cooper and David Manlove</p> <ul> <li>The paper is located at: <a href="https://arxiv.org/abs/2001.10875">https://arxiv.org/abs/2001.10875</a></li> <li>The software is located at: <a href="https://zenodo.org/record/3630383">https://zenodo.org/record/3630383</a></li> <li>The data is located at: <a href="https://zenodo.org/record/3630349">https://zenodo.org/record/3630349</a></li> </ul> <p>See the README for more information.</p>
Supporting data for "The methylome of Biomphalaria glabrata and other mollusks: enduring modification of epigenetic landscape and phenotypic traits by a new DNA methylation inhibitor"
<p>Methylome of the fresh water snail <em>Biomphalaria glabrata</em>. DNA was extracted from the feet of 10 individuals of <em>B. glabrata</em> originally isolated from Brazil. These snails have been cultivated in the laboratory since 1960. Tissue were grinded at 4°C and incubated in 1 ml volume of lysis buffer (20 mM TRIS pH 8; 1 mM EDTA; 100 mM NaCl; 0.5% SDS), with 0.3 mg of proteinase K at 55°C for 1 night. Afterwards, lysate was purified with phenol-chloroform and DNA was isopropanol precipitated. The extracted DNA (around 138ng/µL) was poled in equivalent amounts and Whole Genome Bisulfite Sequencing was done by GATC-biotech (www.gatc-biotech.com). The principle of this treatment is to convert non-methylated cytosines of gDNA into deoxy-uracil, whereas methylated cytosines remain intact. WGBS was done according to the Lister protocol (sequence 2 forward strands only). The reference genome (Biomphalaria-glabrata-BB02_SCAFFOLDS_BglaB1.fa) and annotation (Biomphalaria-glabrata-BB02_BASEFEATURES_BglaB1.3.gff3) used in this project are available on VectorBase (https://www.vectorbase.org/). To align our short reads, we chose to use two specific bisulfite mapping tools, BSMAP 1.0.0 (https://code.google.com/p/bsmap/) and Bismark 0.10.2 (www.bioinformatics.babraham.ac.uk /projects/bismark/), to compare their efficiency and convenience to finally work with the more suitable one on our datasets. IGV (Interactive Genomics Viewer, https://www.broadinstitute.org/igv/) was used to visualized final alignments.<br> BSMAP performed better than Bismark and was used for downstream analyses. Without default parameters alignement efficiency for BSMAP is 47.1%, allowing for 2 mismatches increases it to 55.6%. Methylation occurs predominantly in CpGs. (C methylated in CpG context: 12.4%, C methylated in CHG context: 0.5%, C methylated in CHH context: 0.5%) The major part of CpG sites, 95.7% were unmethylated, of the remaining 4.3% of CpG sites around 3.8% had low methylation, and 0.5% were completely methylated. Methylation is of the mosaic type. Methylation is relatively low with 1.2% of total cytosines. Our analyses suggested that conserved genes and genes with stable expression are localized in high methylated regions of the genome. Finally, we see that repetitive sequences were predominantly situated in low methylated regions of <em>B. glabrata</em>. </p> <p>Wiggle files were generated for CpG pairs only.</p> <p>Produced at IHPE (http://ihpe.univ-perp.fr/)</p>
Supporting Data Sets for "New Constraints on the Lunar Optical Space Weathering Rate"
<p>Data Sets supporting "New Constraints on the Lunar Optical Space Weathering Rate" submitted to Geophysical Research Letter on 12/18/2020. See Supporting Information (link TBD).</p>
DATA ANALYSIS - SARS-COV-2 ( Del69-70 VARIANT ) – NEW UK MUTANTS
<p>The data for S - genome sequence analysis known as Del69-70 is under variant of concern ( VOC ) . It is also termed as variant of investigation ( VUI ) . The data for VUI is statistically analysed by datewise and regionwise . The software used for data analysis is CURVE FINDER V.1.4 . The reproducibility of correlation and standard error is reported here for analysis of scattered data an attempt to study the Rational Fit and Harris Fit .</p>
Quantitative Content Analysis Data for Hand Labeling Road Surface Conditions in New York State Department of Transportation Camera Images
<p><strong>Foundational Codebook and Data: </strong></p> <p>Traffic camera images from the New York State Department of Transportation (511ny.org) are used to create a hand-labeled dataset of images classified into to one of six road surface conditions: 1) severe snow, 2) snow, 3) wet, 4) dry, 5) poor visibility, or 6) obstructed. Six labelers (authors Sutter, Wirz, Przybylo, Cains, Radford, and Evans) went through a series of four labeling trials where reliability across all six labelers were assessed using the Krippendorff’s alpha (KA) metric (Krippendorff, 2007). The online tool by Dr. Freelon (Freelon, 2013; Freelon, 2010) was used to calculate reliability metrics after each trial, and the group achieved inter-coder reliability with KA of 0.888 on the 4th trial. This process is known as quantitative content analysis, and three pieces of data used in this process are shared, including: 1) a PDF of the codebook which serves as a set of rules for labeling images, 2) images from each of the four labeling trials, including the use of New York State Mesonet weather observation data (Brotzge et al., 2020), and 3) an Excel spreadsheet including the calculated inter-coder reliability (ICR) metrics and other summaries used to asses reliability after each trial. The data are included in NYSDOT_quantitative_content_analysis.zip.</p> <p>The broader purpose of this work is that the six human labelers, after achieving inter-coder reliability, can then label large sets of images independently, each contributing to the creation of larger labeled dataset used for training supervised machine learning models to predict road surface conditions from camera images. The xCITE lab (xCITE, 2023) is used to store camera images from 511ny.org, and the lab provides computing resources for training machine learning models.</p> <p><strong>Obstructed Class Variation: </strong></p> <p>There are many applications for labeling roadside camera images, and as a variation of the foundational codebook, an addendum codebook provides another version of labeling the obstructed class. Specifically, this variation prioritizes labeling an image as “obstructed” only in extreme circumstances where there is a camera- or image- specific problem that prevents the assessment of any road surfaces. For labelers who want to use this version of the obstructed class (in this document) and also the other five weather-related classes (in the foundational codebook), the guidance is to use both documents in tandem, making sure to use the obstructed rules/definitions in this document while disregarding the obstructed rules/definitions in the foundational codebook. Alternatively, this codebook may be used alone in applications where the goal is to solely classify obstructed vs not obstructed. To ensure reliability and quality of this variation, quantitative content analysis was conducted on this addendum codebook, just as it was for the foundational codebook. Two labelers were tested with a sample of 30 images and achieved inter-coder reliability with Krippendorff's Alpha of 0.934 after one trial. The data, including the addendum codebook and labeling trial data (images and results) are included in ObstructedVariation_quantitative_content_analysis.zip.</p> <p>This material is based upon work supported by the U.S. National Science Foundation under Grant No. RISE-2019758.</p>
Supporting Data: ontophylo: Reconstructing the evolutionary dynamics of phenomes using new ontology-informed phylogenetic methods
<p>This dataset contains all scripts and data for reproducing the analyses of the paper. The README files contain additional information.</p>
A new repository of electrical resistivity tomography and ground penetrating radar data from summer 2022 near Ny-Ålesund, Svalbard.
<p>We present the geophysical data set acquired in summer 2022 close to Ny-Ålesund (Western Svalbard, Brøggerhalvøya peninsula, Norway) as part of the project ICEtoFLUX (MUR/PRA2021 project-0027). The data set is composed of Electrical Resistivity Tomography (ERT) and GroundPenetrating Radar (GPR) surveys, which are well-known geophysical techniques for the characterization of glacial and hydrological processes and features. 18 ERT profiles and 10 GPR lines were acquired, for a total surveyed length of 9.3 km. The data have been organized in a consistent repository that includes both raw and processed (filtered) data. Some representative examples of 2D models of the subsurface are provided, that is, 2D sections of electrical resistivity (from ERT) and 2D radargrams (from GPR). These examples can support the identification of the active layer and the occurrence of spatial variation of soil conditions at depth. The aim of the investigation is to characterize the role of groundwater flow in correspondence of the active layer as well as through and/or below the permafrost. The data set is of major relevance because scant attention has been paid to the publication of geophysical data from the Ny-Ålesund area so far. Moreover, these geophysical data can foster multidisciplinary scientific collaborations in the fields of hydrology, glaciology, climate, geology, geomorphology, etc. To a large extent, the data set can provide new insight into the hydrological dynamics and polar and climate changes studies on the Ny-Ålesund area. </p>
Fracture Data Supporting: 'The 2024 Mw4.8 New Jersey Intraplate Earthquake: Preferential Rupture of an Immature Fault in Frictionally Unstable Basement Rocks'
<p>The spreadsheets contain fracture and paleoslip surface datasets measured across the epicentral region of the April 5, 2024 Mw4.8 New Jersey earthquake. Datasets contain coordinates of outcrops, strike, dip, and trend/plunge or rake (where slickenlines are observed).</p>
Supporting Data for "Impacts of Antarctic ice mass loss on New Zealand climate"
<p>Contains the model output necessary to reproduce the results of "Impacts of Antarctic ice mass loss on New Zealand climate" by Andrew G. Pauling, Inga J. Smith, Jeff K. Ridley, T. Martin, M. Thomas and D. P. Stevens. Submitted for publication to Geophysical Research Letters.</p> <p>Please use the "getdata.sh" script in the Github repository here: LINK to download and extract the data into the correct location for the notebooks to reproduce the results of the paper.</p>
1.3A (Ge337) calibration data for new Ge115 monochromator installed on Echidna Neutron Powder Diffraction Instrument
<p>In early October 2024 the Echidna neutron powder instrument located at the OPAL reactor, ANSTO, installed a new monochromator with Ge115 cut. The present calibration data were collected shortly afterwards from a standard LaB6 sample in a 6mm diameter Vanadium can. The instrument was set to 140 degrees takeoff angle and monochromator angle 85.08 degrees, corresponding to the Ge337 reflection. Raw data in NeXus format are contained in <strong>ECH0034261.nx.hdf</strong>. These data were corrected for variable detector response using the information in <strong>eff_2024-10-06.cif</strong> and pixel vertical positions adjusted according to the table in <strong>vertical_offsets_2024-10-06.txt. </strong>Deviations from the ideal detector 1.25 degree angular spacing were applied using <strong>echidna-Apr2018.ang</strong>. The detector response was then recorrected based on overlapping measurements using the algorithm described in <a href="https://doi.org/10.1107/S1600576718014048">Avdeev and Hester (2018)</a> resulting in a 1D pattern suitable for fitting wavelength and peak shapes. This 1D pattern is provided here as a plain table (<strong>ECH0034261_LaB6.xyd</strong>) and as a pdCIF file (<strong>ECH0034261_LaB6.cif</strong>) including metadata on data collection and reduction. Details of data reduction are described in the above paper, and the data reduction routines used are included in the <a href="https://github.com/Gumtree/Echidna_scripts">Gumtree package as python code</a>.</p>
Publication text: code, data, and new measures
<p>This Zenodo page describes data collection, processing, and different open access data files related to the text of scientific publications from OpenAlex. If you use the code or data, please cite the following paper: </p> <p>Sam Arts, Nicola Melluso, Reinhilde Veugelers; Beyond Citations: Measuring Novel Scientific Ideas and their Impact in Publication Text. <em>The Review of Economics and Statistics</em> 2025; 1–33 doi: <a href="https://doi.org/10.1162/rest_a_01561" target="_blank" rel="noopener">https://doi.org/10.1162/rest_a_01561</a></p> <p> </p>
amel-github/sars-ani: Releasing new data fields in SARS-ANI dataset
<p>2022-06-20 - Release v1.1</p> <p>The original SARS-ANI dataset displayed common and scientific names of the animal host as found in the information source and/or inferred from the literature or expert knowledge.<br> Misspelled animal names and errors in taxonomy can lead to incorrect scientific conclusions and poor policy design. Moreover, harmonized host names can aid integrating other datasets (e.g. data on host biological traits, geographic distribution, or association with other pathogens).<br> Therefore, for each event, we programmatically performed taxonomic validation of the animal host name, using the R package taxize (Chamberlain et al. 2013). For more information on our validation process, see the R script <strong>sars_ani_validation.R.</strong></p> <p>Version 1.1. contains seven fields related to the identification of the animal host:</p> <ul> <li> <p>host_com_orig: Most specific designation of the animal host provided by the source(s), in English.</p> </li> <li> <p>host_sci_orig: Scientific name of the animal host as mentioned in the source(s) (scientific names are harmonized so that only the first letter of the genus is capitalized).</p> </li> <li> <p>host_com_res: Common name of the animal host, harmonized against the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/">NCBI</a>) taxonomic backbone.</p> </li> <li> <p>host_sci_res: Scientific name of the animal host (resolved to species or subspecies level), harmonized against the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/">NCBI</a>) taxonomic backbone.</p> </li> <li> <p>host_colloq: The colloquial name of the host, i.e. the name commonly used to identify the animal in non-specialist language (e.g. "tiger" for "Sumatran tiger").</p> </li> <li> <p>host_sci_spec_res: The scientific name of the host resolved to the species level.</p> </li> <li> <p>family: Animal family of the animal host.</p> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.