Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

8

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

8 results for “Contiguous U.S.”

Learn how ShareScore rates datasets ↗
zenodo40/100

Transformation Rate Maps of Dissolved Organic Carbon in the Contiguous U.S.

<p>We develop two new maps of the dissolved organic carbon (DOC) transformation rate (\(P_r\)) over the contiguous United States. Those maps are derived by combining the USGS riverine DOC observations, soil organic carbon (SOC) data from two sources&mdash;HWSD v1.2 and SoilGrids 2.0, and the watershed characteristics from two existing datasets medium-resolution NHDplus and ScienceBase, and state-of-the-art machine learning techniques.&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

Wind and solar capacity factor time series by year and grid cell over the contiguous U.S.

<p>This data set contains hourly capacity factor time series of wind and solar resources over the contiguous U.S.</p> <p>&nbsp;</p> <p>The included time series cover four individual years and 2,586 grid cells. The years range from 2016 to 2019. The grid cells correspond to the grid cells of the NASA&#39;s MERRA-2 reanalysis data set into which the contiguous U.S. is subdivided. The grid cells have a spatial resolution of 0.5&deg; latitude x 0.625&deg; longitude with dimensions ranging from about 55 km x 45 km to 55 km x 62 km.</p> <p>&nbsp;</p> <p>This data set is used to generate the results of the following journal article:</p> <p>Enrico G. A. Antonini, Tyler H. Ruggles, David J. Farnham, Ken Caldeira, &quot;The quantity-quality transition in the value of expanding wind and solar power generation&quot;, iScience 25 (4), 104140, 2022.</p> <p>&nbsp;</p> <p>Code and instructions required to reproduce the results reported in the above paper are available in the GitHub repositories at <a href="https://github.com/eantonini/Distributed_wind_and_solar_generation">https://github.com/eantonini/Distributed_wind_and_solar_generation</a> and <a href="https://github.com/carnegie/MEM_public/tree/Antonini_et_al_2022">https://github.com/carnegie/MEM_public/tree/Antonini_et_al_2022</a>.</p>

opencc-by-4.0Apr 2022View details →
dryad36/100

Data from: Winners and losers from climate change: An analysis of tree growth and survival responses to temperature and precipitation for roughly 150 species across the contiguous U.S.

Open the record for dataset details and reuse information.

publicNov 2024View details →
zenodo32/100

Using knowledge-guided machine learning to assess patterns of areal change in waterbodies across the contiguous U.S.: Data

<p>Data used for generating figures in the knowledge-guided machine learning manuscript and supplemental info by Wander et al.</p><p>This repository contains three folders:</p><ul><li><strong>Results: </strong>csv file with the final KGML groups for each waterbody. Latitude, longitude, RealSAT (Khandelwal et al., 2022) waterbody id, and HydroLAKES (Messager et al., 2016) waterbody id are provided.</li><li><strong>Code_data</strong>: data used for generating figures in the knowledge-guided machine learning manuscript and supplemental info by Wander et al.</li><li><strong>prism</strong>: data from PRISM dataset (Matsuura and Willmott) used for preliminary driver analysis in Wander et al.</li></ul><p>Khandelwal, A.; Karpatne, A.; Ravirathinam, P.; Ghosh, R.; Wei, Z.; Dugan, H.; Hanson, P. C.; Kumar, V. ReaLSAT, a global dataset of reservoir and lake surface area variations. <i>Sci. Data</i> <strong>2022</strong>, <i>9</i>, 356. <a href="https://doi.org/10.1038/s41597-022-01449-5">https://doi.org/10.1038/s41597-022-01449-5</a></p><p>Messager, M. L.; Lehner, B.; Grill, G.; Nedeva, I.; Schmitt, O. Estimating the volume and age of water stored in global lakes using a geo-statistical approach. <i>Nat. Commun.</i> <strong>2016</strong>, <i>7</i>, 1–11. <a href="https://doi.org/10.1038/ncomms13603">https://doi.org/10.1038/ncomms13603</a><br><br>Willmott, C. J.; Matsuura, K. Terrestrial Air Temperature and Precipitation: 1900-2014 Gridded Monthly Time Series: NOAA Physical Sciences Laboratory Terrestrial Air Temperature and Precipitation: 1900-2014 Gridded Monthly Time Series, <strong>2015</strong>. <a href="https://psl.noaa.gov/data/gridded/data.UDel_AirT_Precip.html">https://psl.noaa.gov/data/gridded/data.UDel_AirT_Precip.html</a>. (Accessed Jan 2022).</p><p>Danielson, J. J.; Gesch, D. B. TEMIS -- GMTED2010 Elevation Data at Different Resolutions, <strong>2011</strong>. <a href="https://www.usgs.gov/centers/eros/science/usgs-eros-archive-digital-elevation-global-multi-resolution-terrain-elevation">https://www.usgs.gov/centers/eros/science/usgs-eros-archive-digital-elevation-global-multi-resolution-terrain-elevation</a>. (Accessed Jan 2022).</p>

opencc-by-4.0May 2023View details →
zenodo24/100

Stacked multiband receiver functions beneath contiguous U.S. and preferred inverted parameters

<p>The preferred 12 parameters and their uncertainties (Table S2). Grouped and stacked receiver functions from the contiguous U.S. See the readme file for detail format.</p>

opencc-by-4.0Dec 2022View details →
nasa24/100

Daily and Annual PM2.5, O3, and NO2 Concentrations at ZIP Codes for the Contiguous U.S., 2000-2016, v1.0

The Daily and Annual PM2.5, O3, and NO2 Concentrations at ZIP Codes for the Contiguous U.S., 2000-2016, v1.0 data set contains daily and annual concentration predictions for Fine Particulate Matter (PM2.5), Ozone (O3), and Nitrogen Dioxide (NO2) pollutants at ZIP Code-level for the years 2000 to 2016. Ensemble predictions of three machine-learning models were implemented (Random Forest, Gradient Boosting, and Neural Network) to estimate the daily PM2.5, O3, and NO2 at the centroids of 1km x 1km grid cells across the contiguous U.S. for 2000 to 2016. The predictors included air monitoring data, satellite aerosol optical depth, meteorological conditions, chemical transport model simulations, and land-use variables. The ensemble models demonstrated excellent predictive performance with 10-fold cross-validated R-squared values of 0.86 for PM2.5, 0.86 for O3, and 0.79 for NO2. These high-resolution, well-validated predictions allow for estimates of ZIP Code-level pollution concentrations with a high degree of accuracy. For general ZIP Codes with polygon representations, pollution levels were estimated by averaging the predictions of grid cells whose centroids lie inside the polygon of that ZIP Code; for other ZIP Codes such as Post Offices or large volume single customers, they were treated as a single point and predicted their pollution levels by assigning the predictions using the nearest grid cell. The polygon shapes and points with latitudes and longitudes for ZIP Codes were obtained from Esri and the U.S. ZIP Code Database and were updated annually. The data include about 31,000 general ZIP Codes with polygon representations, and about 10,000 ZIP Codes as single points. The aggregated ZIP Code-level, daily predictions are applicable in research such as environmental epidemiology, environmental justice, health equity, and political science, by linking with ZIP Code-level demographic and medical data sets, including national inpatient care records, medical claims data, census data, U.S. Census Bureau American CommUnity Survey (ACS), and Area Deprivation Index (ADI). The data are particularly useful for studies on rural populations who are under-represented due to the lack of air monitoring sites in rural areas. Compared with the 1km grid data, the ZIP Code-level predictions are much smaller in size and are manageable in personal computing environments. This greatly improves the inclusion of scientists in different fields by lowering the key barrier to participation in air pollution research. The Units are ug/m^3 for PM2.5 and ppb for O3 and NO2.

restrictednotspecifiedApr 2025View details →
nasa24/100

Annual Mean PM2.5 Components (EC, NH4, NO3, OC, SO4) 50m Urban and 1km Non-Urban Area Grids for Contiguous U.S., 2000-2019 v1

The Annual Mean PM2.5 Components (EC, NH4, NO3, OC, SO4) 50m Urban and 1km Non-Urban Area Grids for Contiguous U.S., 2000-2019, v1 data set contains annual predictions of the chemical concentrations at a hyper resolution (50m x 50m grid cells) in urban areas and at a high resolution (1km x 1km grid cells) in non-urban areas for the years 2000 to 2019. Particulate matter with an aerodynamic diameter less than 2.5 �m (PM2.5) increases mortality and morbidity. PM2.5 is composed of a mixture of chemical components that vary across space and time. Due to limited hyperlocal data availability, less is known about health risks of PM2.5 components, their U.S.-wide exposure disparities, or which species are driving the biggest intra-urban changes in PM2.5 mass. The national super-learned models were developed across the U.S. for hyperlocal estimation of annual mean elemental carbon, ammonium, nitrate, organic carbon, and sulfate concentrations across 3,535 urban areas at a 50m spatial resolution, and at a 1km resolution for non-urban areas from 2000 to 2019. Using Machine-Learning models (ML), combined with either a Generalized Additive Model (GAM) Ensemble Geographically-Weighted-Averaging (GAM-ENWA) or Super-Learning (SL) and approximately 82 billion predictions across 20 years, hyperlocal super-learned PM2.5 components are now available for further research. The overall R-squared values of 10-fold cross validated models ranged from 0.910 to 0.970 on the training sets for these components, while on the test sets the R-squared values ranged from 0.860 to 0.960. Remarkable spatiotemporal intra-urban and inter-urban variabilities were found in PM2.5 components. The Coordinate Reference System (CRS) for predictions is the World Geodetic System 1984 (WGS84) and the Units for the PM2.5 Components are �g/m^3. The data are provided in RDS tabular format, a file format native to the R programming language, but can also be opened by other languages such as Python.

restrictednotspecifiedApr 2025View details →
nasa20/100

Annual Mean PM2.5 Components Trace Elements (TEs) 50m Urban and 1km Non-Urban Area Grids for Contiguous U.S., 2000-2019, v1

The Annual Mean PM2.5 Components Trace Elements (TEs) 50m Urban and 1km Non-Urban Area Grids for Contiguous U.S., 2000-2019, v1 data set contains annual predictions of trace elements concentrations at a hyper resolution (50m x 50m grid cells) in urban areas and a high resolution (1km x 1km grid cells) in non-urban areas, for the years 2000 to 2019. Particulate matter with an aerodynamic diameter of less than 2.5 �m (PM2.5) is a human silent killer of millions worldwide, and contains many trace elements (TEs). Understanding the relative toxicity is largely limited by the lack of data. In this work, ensembles of machine learning models were used to generate approximately 163 billion predictions estimating annual mean PM2.5 TEs, namely Bromine (Br), Calcium (Ca), Copper (Cu), Iron (Fe), Potassium (K), Nickel (Ni), Lead (Pb), Silicon (Si), Vanadium (V), and Zinc (Zn). The monitored data from approximately 600 locations were integrated with more than 160 predictors, such as time and location, satellite observations, composite predictors, meteorological covariates, and many novel land use variables using several machine learning algorithms and ensemble methods. Multiple machine-learning models were developed covering urban areas and non-urban areas. Their predictions were then ensembled using either a Generalized Additive Model (GAM) Ensemble Geographically-Weighted-Averaging (GAM-ENWA), or Super-Learners. The overall best model R-squared values for the test sets ranged from 0.79 for Copper to 0.88 for Zinc in non-urban areas. In urban areas, the R-squared model values ranged from 0.80 for Copper to 0.88 for Zinc. The Coordinate Reference System (CRS) used in the predictions is the World Geodetic System 1984 (WGS84) and the Units for the PM2.5 Components TEs are ng/m^3. The data are provided in RDS tabular format, a file format native to the R programming language, but can also be opened by other languages such as Python.

restrictednotspecifiedApr 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record