Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Data from: Effects of input data sources on species distribution model predictions across species with different distributional ranges
<p>Species distribution models (SDMs) are a popular tool in theoretical and quantitative ecology, and constitute the most widely used modelling framework in global change science and biodiversity conservation. As main data sources, SDMs require georeferenced biodiversity observations as a response or dependent variable (e.g. species occurrence, species richness, etc) and geographic layers of environmental information as predictors or independent variables (e.g. climate, land cover, vegetation indices derived from remote sensing, etc). However, although SDMs have become one of the most important quantitative tools for addressing regular and timely biodiversity assessments worldwide, these techniques are still subject to different sources of uncertainty that have been unequally assessed. Thus, despite uncertainty related to niche-based or distribution-based models has been addressed at different stages in the modelling process, an analysis of the effect of uncertainty coming from alternative data sources on the predictive ability of SDMs is still limited.</p> <p>Citizen-collected species occurrence data (e.g. eBird) are often used for fitting SDMs when data from standardized and expert-supported surveys (e.g. Atlases) are unavailable. On the other hand, macroclimate variables are much more commonly used as predictors in SDMs than other sources of information coming from remote sensing data. We assessed the effects of using different data sources (in both response and predictor variables) on SDM performance across a wide range of bird species with contrasting distributional ranges in the Iberian Peninsula (Portugal and Spain). To do that, a SDM ensemble-forecasting approach was implemented by using bird data from two different data sources: the semi-structured eBird project and standardized Atlases. We fitted SDMs with three predictor types: macroclimate, remotely sensed ecosystem functional attributes (EFAs) from vegetation indices, and their combination. Species were grouped in four range size classes. We also used different evaluation metrics to better assess the uncertainty of model predictions. We then applied generalized linear mixed-effects models to test the effect on model performance of input data source across distributional range sizes while accounting for different accuracy metrics. Pairwise comparisons between range projections were used to assess their spatial similarity.</p> <p>Our models demonstrated the usefulness and complementarity of different input data sources when modelling species distribution across different distributional ranges. Citizen science and remote sensing data contribute to update the knowledge of the distribution of the most threatened bird species by increasing the model accuracy. These findings highlight the need to integrate different data sources to improve the model predictions at regional scale. Our framework also underlines that model uncertainty should be examined more exhaustively at early stages of the modelling process.</p> <p>To perfom and replicate this study, this dataset provides all needed files (as tables) to fit SDMs: i) the Iberian bird species occurrences at 10km UTM square as a response or dependent variable; ii) the geographic layers of environmental information at 10km UTM square for the Iberian Peninsula as predictors or independent variables, such as climate data, ecosystem functioning attributes (EFAs) and the combined climate and EFA data. The dataset is provided by four <em>*.csv</em> files named as:</p> <p><em>1) The_Iberian_bird_species_occurrences_dataset_10km.csv</em></p> <p><em>2) CHELSA_bioclimate_variables_IP10km.csv</em></p> <p><em>3) MODIS_EVI-based_EFAs_IP10km.csv</em></p> <p><em>4) Combined_bioclimate_EFA_dataset_IP10km.csv</em></p> <p>For a more detailed description of the main dataset and each of these subdatasets, please refer to the attached README file.</p> <p><strong>Keywords:</strong> bird atlas, eBird data, ecosystem functional attributes (EFAs), Iberian Peninsula, IUCN categories, Model accuracy, MODIS EVI, narrow-ranged species, remote sensing, species distribution models (SDMs), widespread species</p>
Finding Efficient Trade-offs in Multi-Fidelity Response Surface Modeling: Generated data files and figures
<p>All data files and figures generated for the paper "Finding Efficient Trade-offs in Multi-Fidelity Response Surface Modeling".</p> <p>The code used to generate this is archived at <a href="https://doi.org/10.5281/zenodo.6123254">zenodo.org/record/6123254</a></p>
Evaluation data for "Global, high-resolution, reduced-complexity air quality modeling for PM2.5 using InMAP (Intervention Model for Air Pollution)"
<p>This zip file contains data for performing Global InMAP model runs and evaluations. To the extent that any of the data is covered by third party licenses, it is the responsibility of the user to follow the terms of those licenses. A description of the contents of this directory is below:</p> <p>measurements.csv<br> Vetted global dataset of ground-level annual-average measurements of total PM2.5 and species (pNO3, pSO4, pNH4) compiled from monitoring networks, used for model performance evaluation. Data sources are: World Health Organization (Global), European Environment Agency (Europe), National Air Pollution Surveillance Program (Canada), Environmental Protection Agency (United States of America), Central Pollution Control Board (India), Australian Government State of the Environment (Australia), and Acid Deposition Monitoring Network In East Asia (EANET) (East Asia).</p> <p>population directory<br> Population count data is from the Gridded Population of The World (v4.10) projected to year 2020. The data is in 15x15 arcminute grids, except for in grid cells where the population is above 80,000, where the population data is 30x30 arcseconds.</p> <p>GlobalInMAPData_v1.ncf<br> Regular-grid Global InMAP input data for the year 2005 for use as the "InMAPData" variable in the InMAP configuration file. It was created from GEOS-Chem v.11-01 simulation outputs with the 'inmap preproc' command.</p> <p>global_inmap_004x003_v1.1.0.gob<br> Global InMAP variable grid resolution input data for coords for year 2016 for use as the "VariableGridData" variable in the InMAP configuration file. It was created with the 'inmap grid' command using GlobalInMAPData_v1.ncf and population.shp.</p> <p>2016_emissions directory<br> Total PM2.5 and precursor emissions to arrive at total PM2.5 concentrations from Global InMAP. Units for polygonized emissions inputs (shapefiles) are short (US) tons/yr, and units for gridded emissions inputs (NetCDF files) are kg/yr.</p> <p>global_emission_changes directory<br> nh3.nc, nox.nc, and sox.nc are gridded emissions for changes in inorganic precursors for comparing Global InMAP and GEOS-Chem. Units are kg/yr. NH4-gc.nc, NIT-gc.nc, and SO4-gc.nc are results for changes in concentrations arising from these changes in emissions for 3 months, 1 month, and 2 months.</p> <p>usa_emission_changes directory<br> Emissions for comparing Global InMAP and US InMAP (described in Tessum et al., 2017).<br> Emissions are derived using the United States National Emissions Inventory (NEI) 2014v.1, processed exactly as in Thakrar et al., 2020.<br> Emissions are coal-powered electricity generation (NEI Source Classification Code: 10100212) and gasoline passenger vehicles (NEI Source Classification Code: 2201210080).<br> Units are ug/s.</p> <p>Tessum, C.W.; Hill, J.D.; Marshall, J.D. InMAP: A model for air pollution interventions. PloS One 2017, 12 (4) e0176131.<br> Thakrar, S.K.; Balasubramanian, S.; Adams, P.J.; Azevedo, I.M.; Muller, N.Z.; Pandis, S.N.; Polasky, S.; Pope III, C.A.; Robinson, A.L.; Apte, J.S.; Tessum, C.W.; Marshall, J.D.; Hill; J.D. Reducing mortality from air pollution in the United States by targeting specific emission sources. Environmental Science & Technology Letters 2020, 7(9), pp.639-645.<br> Gridded Population of the World, Version 4 (GPWv4): National Identifier Grid. Palisades, NY: NASA Socioeconomic Data and Applications Center (SEDAC). http://dx.doi.org/10.7927/H41V5BX1.</p>
Data for the article "Solution of the Thirring model in thimble regularization"
<p>Data set for the paper "Solution of the Thirring model in thimble regularization", arXiv:2109.02511 [hep-lat]. It includes the data required to generate the figures in the article.</p> <p> </p>
Raw Data for the article: Atrial Septostomy for Left Ventricular Unloading During Extracorporeal Membrane Oxygenation for Cardiogenic Shock: Animal Model
<p><strong>Objectives: </strong>The aim of this study was to quantify and understand the unloading effect of percutaneous balloon atrial septostomy (BAS) in acute cardiogenic shock (CS) treated with venoarterial (VA) extracorporeal membranous oxygenation (ECMO).</p> <p><strong>Background: </strong>In CS treated with VA ECMO, increased left ventricular (LV) afterload is observed that commonly interferes with myocardial recovery or even promotes further LV deterioration. Several techniques for LV unloading exist, but the optimal strategy and the actual extent of such procedures have not been fully disclosed.</p> <p><strong>Methods: </strong>In a porcine model (n = 11; weight 56 kg [53-58 kg]), CS was induced by coronary artery balloon occlusion (57 minutes [53-64 minutes]). Then, a step-up VA ECMO protocol (40-80 mL/kg/min) was run before and after percutaneous BAS was performed. LV pressure-volume loops and multiple hemoglobin saturation data were evaluated. The Wilcoxon rank sum test was used to assess individual variable differences.</p> <p><strong>Results: </strong>Immediately after BAS while on VA ECMO support, LV work decreased significantly: pressure-volume area, end-diastolic pressure, and stroke volume to ∼78% and end-systolic pressure to ∼86%, while superior vena cava and tissue oximetry did not change. During elevating VA ECMO support (40-80 mL/kg/min) with BAS vs without BAS, we observed 1) significantly less mechanical work increase (122% vs 172%); 2) no end-diastolic volume increase (100% vs 111%); and 3) a considerable increase in end-systolic pressure (134% vs 144%).</p> <p><strong>Conclusions: </strong>In acute CS supported by VA ECMO, atrial septostomy is an effective LV unloading tool. LV pressure is a key component of LV work load, so whenever LV work reduction is a priority, arterial pressure should carefully be titrated low while maintaining organ perfusion.</p>
Raw Data for the article: Patient-Specific Analysis of Ascending Thoracic Aortic Aneurysm with the Living Heart Human Model
<p>In ascending thoracic aortic aneurysms (ATAAs), aneurysm kinematics are driven by ventricular traction occurring every heartbeat, increasing the stress level of dilated aortic wall. Aortic elongation due to heart motion and aortic length are emerging as potential indicators of adverse events in ATAAs; however, simulation of ATAA that takes into account the cardiac mechanics is technically challenging. The objective of this study was to adapt the realistic Living Heart Human Model (LHHM) to the anatomy and physiology of a patient with ATAA to assess the role of cardiac motion on aortic wall stress distribution. Patient-specific segmentation and material parameter estimation were done using preoperative computed tomography angiography (CTA) and ex vivo biaxial testing of the harvested tissue collected during surgery. The lumped-parameter model of systemic circulation implemented in the LHHM was refined using clinical and echocardiographic data. The results showed that the longitudinal stress was highest in the major curvature of the aneurysm, with specific aortic quadrants having stress levels change from tensile to compressive in a transmural direction. This study revealed the key role of heart motion that stretches the aortic root and increases ATAA wall tension. The ATAA LHHM is a realistic cardiovascular platform where patient-specific information can be easily integrated to assess the aneurysm biomechanics and potentially support the clinical management of patients with ATAAs.</p>
Antarctic surface climate and surface mass balance in the Community Earth System Model version 2 (1850-2100) - AWS data
<p>This Antarctica AWS temperature and wind speed dataset was compiled by Alexandra Gossart and Niels Souverijns (<a href="https://doi.org/10.1175/JCLI-D-19-0030.1">https://doi.org/10.1175/JCLI-D-19-0030.1</a>).</p>
4D-Var data assimilation experiment of the Lorenz 96 model using an adjoint model of a neural network surrogate model
<p>These data are the output of the 4D-Var data assimilation experiment of the Lorenz96 model using an adjoint model of a neural network surrogate model.<br> The details are described in Nishizawa (2022).<br> </p>
Data publication for "First-principles derivation and properties of density-functional average-atom models"
<p>Data for the pre-print "First-principles derivation and properties of density-functional average-atom models", https://arxiv.org/abs/2103.09928.</p> <p>Each data folder is named according to the corresponding figure in the paper. For any questions, please contact the authors.</p>
Neodymium isotopes as a paleo-water mass tracer: A model-data reassessment: Model Output Data
<p>This dataset contains model output for the simulations presented in <em>"Neodymium isotopes as a paleo-water mass tracer: A model-data reassessment, Quaternary Science Reviews 279 (2022), 107404".</em></p> <p>The NetCDF4 files contain the following variables:</p> <p>3D fields:</p> <ul> <li>Potential Temperature</li> <li>Salinity</li> <li>North Atlantic dye tracer</li> <li>Epsilon Nd</li> <li>Nd concentration</li> </ul> <p>2D field:</p> <ul> <li>AMOC stream function</li> </ul>
Data for RSTA-2021-0255 - Modelling attenuation of irregular wave fields by artificial ice floes in the laboratory
<p>Data are from two independent experimental campaigns in 3D wave basins, which were conducted to investigate the interaction between surface waves and thin floating plates, mimicking the interaction between ocean waves and sea ice floes. </p> <p>One experiment tested irregular waves, including directional spreading, with two peak steepness values comparable to storm and polar cyclone conditions, interacting with a single square polypropylene floe, with length close to the dominant wavelength (scattering regime) -- data file Single_Plate.zip.</p> <p>The second experiment used arrays of wooden floes with low- and high-concentrations to test wave-floe interaction at a range of peak periods and peak wave steepness spanning from gentle to storm-like conditions -- data file MultipleDisks.zip.</p> <p>Both data files contains raw data, experimental matrix and experimental set.</p> <p>Data are fully described and analysed in a contribution to "Theory, modelling and observations of marginal ice zone dynamics: Multidisciplinary perspectives and outlooks" a special theme issue for Philosophical Transactions A: </p> <p>Toffoli A., Pitt J. P. A., Alberello A., Bennetts L.G., 2022. Modelling attenuation of irregular wave fields by artificial ice floes in the laboratory. Philos. T. Roy. Soc. A (submitted)</p> <p>Additional information on the data sets can be found in </p> <p>Bennetts LG, Alberello A, Meylan MH, Cavaliere C, Babanin AV, Toffoli A. 2015 An idealised experimental model of ocean surface wave transmission by an ice floe, Ocean Model. 96, 85–92 </p> <p>Bennetts LG, Williams TD. 2015 Water wave transmission by an array of floating discs, Proc. Roy. Soc. A 471, 20140698</p>
Marsquake locations and 1-D seismic models for Mars from InSight data
<p>Data used to draw the figures in the paper 'Marsquake locations and 1-D seismic models for Mars from InSight data'.</p>
Bootstrap methods for quantifying the uncertainty of binding constants in the hard modeling of spectrophotometric titration data
<p>Supporting simulation code and data for the manuscript "Bootstrap methods for quantifying the uncertainty of binding constants in the hard modeling of spectrophotometric titration data"</p>
Data for "International Transport costs: New Findings from modeling additive costs"
<p>Downloading the data will provide you with a .zip file.</p> <p>These data are US imports from 1974 to 2020, by partner, transport mode and product at the 5 digit level in the "Hummels" data and by partner, transport mode, district of entry, district of unlading and product at the 10 digit level for the rest of the data.</p> <p>All the data originally come from the Census Bureau "Foreign Trade" (see https://www.census.gov/foreign-trade/index.html) and belong to the public domain.</p> <p>"Hummels_JEP _data" was downloaded from David Hummels’s website (https://www.krannert.purdue.edu/faculty/hummelsd/research/jep/data.html -- Hummels, David, "Transportation Costs and International Trade in the Second Era of Globalization", <em>Journal of Economic Perspectives</em>, Vol 21, No 3, pp 131-154. It covers 1974 to 2004.</p> <p>The other main files were bought directly from the Census Bureau ("Annual Merchandise Trade files"). They cover 1997-1999 and 2002-2020.</p> <p>Various files are included for the conversion of country codes (dist_cepii.dta, countrycodes_use.txt), product codes (HS2002_SITC2.txt) and to associate quantity units with hts codes (hts... and ..._hts_...).</p>
A reliable algorithm for calculating stoichiometry parameters in the hard modeling of spectrophotometric titration data
<p>Supporting Matlab code and data for the manuscript "A reliable algorithm for calculating stoichiometry parameters in the hard modeling of spectrophotometric titration data"</p>
Storyline data used in the paper "The July 2019 European heatwave in a warmer climate: Storyline scenarios with a coupled model using spectral nudging"
<p>We provide the storyline data (in NetCDF format) used in the paper: “The July 2019 European heatwave in a warmer climate: Storyline scenarios with a coupled model using spectral nudging” published in Journal of Climate. The data is structured in four .tar.gz files (Preindustrial, Present, 2 and 4 K warmer climates) containing all variables used in this each climate. The data from the five ensemble members (E1 to E5) have been included separately in 3-months files.</p> <p>Atmospheric variables (Files are named as: {variable}_E{ensemble member}_{starting month}{year}.nc:</p> <ul> <li> <p>Latent heat flux (ahfl)</p> </li> <li> <p>Sensible heat flux (ahfs)</p> </li> <li> <p>Monthly Global Mean 2m Temperature (GMTT2mMonthly)</p> </li> <li> <p>Maximum 2m Temperature (t2max)</p> </li> <li> <p>Mean 2m Temperature (t2mean)</p> </li> <li> <p>Minimum 2m Temperature (t2max)</p> </li> <li> <p>Soil Wetness (ws)</p> </li> </ul> <p> Only for present climate:</p> <ul> <li> <p>850 hPa Temperature (T850)</p> </li> <li> <p>Total Cloud Cover (TCC)</p> </li> <li> <p>500 hPa Geopotential Height (Z500)</p> </li> </ul> <p>Five layers soil moisture (Only for present climate, Files are named as: From20172019in2017Climatessp370{ensemble member}_{year}{starting month}.01_jsbid.nc) </p> <p>Oceanic variables (from FESOM, Files are named as: {variable}_E{ensemble member}_{year}{starting month}01.nc:</p> <ul> <li> <p>Sea Ice Concentration (SIC)</p> </li> <li> <p>Sea Surface Temperature (SST)</p> </li> </ul> <p><strong>Please, note that FESOM uses an unstructured mesh.</strong></p>
Data for: Modeling of spatial pattern and influencing factors of cultivated land quality based on spatial-temporal big data (PONE-D-21-21084R1)
<p>The quality of cultivated land determines the production capacity of cultivated land and the level of regional development, and also directly affects the food security and ecological safety of the country. This paper starts from the perspective of spatial pattern of cultivated land quality and uses spatial autocorrelation analysis to study the spatial aggregation characteristics and differences of cultivated land quality in Henan Province at the county level scale, and also uses bivariate spatial autocorrelation to analyze the influence of neighboring influences on the quality of cultivated land in the target area. The spatial autoregressive model was used to further analyze the driving factors affecting the quality of cultivated land, and the influence of cultivated land area index was coupled in the process of rating analysis, which was finally used as a basis to propose more precise measures for the protection of cultivated land zoning. The results show that: (1) The quality of cultivated land in Henan Province has a strong spatial correlation (global Moran's I≈0.710) and shows an obvious aggregation pattern in spatial distribution; positive correlation types (high-high and low-low) are concentrated in north-central and western mountainous areas of Henan Province, respectively; negative correlation types are discrete. The negative correlation types are distributed in a discrete manner. (2) The bivariate spatial autocorrelation results show that Slope (Moran's I≈-0.505), Irrigation guarantee rate (IGR, 0.354), Urbanization rate (-0.255), Total agricultural machinery power (TAMP, 0.331) and Pesticide use (0.214) are the main influencing factors. (3) According to the absolute values of the regression coefficients, it can be seen that the magnitude of the influence of different factors on the quality of cultivated land is: Slope (0.089) >IGR (0.025) > Urbanization rate (0.002) > TAMP (0.001) > Pesticide use (1.96e-006). (4) Based on the spatial pattern presented by the spatial autocorrelation results, we proposed corresponding protection zoning measures to provide more scientific reference decisions and technical support for the implementation of refined cultivated land management in Henan Province. </p>
Divergence time estimation using ddRAD data and an isolation-with-migration model applied to water vole populations of Arvicola
<p>Molecular dating methods of population splits are crucial in evolutionary biology, but they present important difficulties due to the complexity of the genealogical relationships of genes and past migrations between populations. Using the double digest restriction-site associated DNA (ddRAD) technique and an isolation-with-migration (IM) model, we studied the evolutionary history of water vole populations of the genus <em>Arvicola</em>, a group of complex evolution with fossorial and semi-aquatic ecotypes. To do this, we first estimated mutation rates of ddRAD loci using a phylogenetic approach. An IM model was then used to estimate split times and other relevant demographic parameters. A set of 300 ddRAD loci that included 85 calibrated loci resulted in good mixing and model convergence. The results showed that the two populations of <em>A. scherman</em> present in the Iberian Peninsula split 34 thousand years ago, during the last glaciation. In addition, the much greater divergence from its sister species, <em>A. amphibius</em>, may help to clarify the controversial taxonomy of the genus. We conclude that this approach, based on ddRAD data and an IM model, is highly useful for analyzing the origin of populations and species.</p>
Data of the protocol paper: Protein Structural Modeling for Electron Microscopy Maps Using VESPER and MAINMAST
<p>The archived file contains data of the test cases used used in the protocol paper. For each protocol, there's a folder containing input files the protocol takes, and a folder for the protocol output files.</p>
Data used in JAMES paper "Sensitivity of the Horizontal Scale of Convective Self‐Aggregation to Sea Surface Temperature in Radiative Convective Equilibrium Experiments Using a Global Nonhydrostatic Model"
<p>Data used in JAMES paper "Sensitivity of the Horizontal Scale of Convective Self‐Aggregation to Sea Surface Temperature in Radiative Convective Equilibrium Experiments Using a Global Nonhydrostatic Model" by Shuhei Matsugishi and Masaki Satoh doi: 10.1029/2021MS002636</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.