Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Supporting data for the Total ELectricity Loads (TELL) model
<p>This is the raw data supporting the Total ELectricity Loads (TELL) model developed by the Integrated Multisector Multiscale Modeling (IM3) project led by Pacific Northwest National Laboratory. Information about the model itself and code to extract and work with this raw data can be found at: https://github.com/IMMM-SFA/tell.</p> <p>There are five core datasets in this package:</p> <p>1) EIA_930 contains data on hourly electricity demand for 68 balancing authorities in the U.S.</p> <p>2) EIA_861 contains an annual summary of the characteristics of the power industry in the U.S.</p> <p>3) Population contains annual estimates of county- and state-level population in the U.S.</p> <p>4) County_Shapefiles contains a set of shapefiles that define county boundaries for counties in the U.S.</p> <p>4) State_Shapefiles contains a set of shapefiles that define state boundaries for states in the U.S.</p>
On the estimation of landslide intensity, hazard and density via data-driven models.
<p>The geographic prediction of landslide occurrence is undertaken by assessing whether a slope may be stable or unstable. In other words, current practices treat slopes where a single landslide occurred in the same way as slopes where many landslides occurred. At the slope scale, this procedure inevitably underestimates the effect of multiple landslides.<br> Here we model the number of landslides per slope instead. Then, thanks to the close relation that the number of failures shows with respect to landslide size, we convert the estimated number of landslides into estimated landslide areas. Ultimately, we also estimate the expected proportion of a slope affected by landslides. This framework is more informative than the stable/unstable paradigm and may help landslide risk mitigation strategies.</p>
Biogeomorphic modeling to assess the resilience of tidal-marsh restoration to sea level rise and sediment supply - Supporting code and data
<p>Code and data to reproduce figures and analyses of the paper:</p> <p>Gourgue, O., van Belzen, J., Schwarz, C., Vandenbruwaene, W., Vanlede, J., Belliard, J.-P., Fagherazzi, S., Bouma, T.J., van de Koppel, J., and Temmerman, S.: Biogeomorphic modeling to assess resilience of tidal marsh restoration to sea level rise and sediment supply, Earth Surf. Dynam., submitted.</p> <p>Standard Python dependencies:</p> <ul> <li>GDAL</li> <li>Geopandas</li> <li>Matplotlib</li> <li>NumPy</li> <li>Rasterio</li> <li>SciPy</li> <li>Seaborn</li> <li>Shapely</li> <li>scikit-learn</li> </ul> <p>Third-party Python dependencies:</p> <ul> <li>Centerline (https://github.com/fitodic/centerline)</li> <li>pputils (https://github.com/pprodano/pputils)</li> <li>pysheds (https://github.com/mdbartos/pysheds)</li> </ul> <p>In-house Python dependencies:</p> <ul> <li>Demeter 1.0.5 (https://doi.org/10.5281/zenodo.5205258)</li> <li>OGTools 1.1 (https://doi.org/10.5281/zenodo.3994952)</li> <li>TidalGeoPro 0.1 (https://doi.org/10.5281/zenodo.5205285)</li> </ul>
Data used for the article "Hybrid intrahour DNI forecast model based on DNI measurements and sky-imaging data"
<p>Data used to obtain the results presented in the article "Hybrid intrahour DNI forecast model based on DNI measurements and sky-imaging data".</p> <ul> <li>CNRS_PROMES_DNI_2020-09-03_2021-01-11.zip : contains DNI measurements taken at PROMES-CNRS laboratory in Odeillo.</li> <li>The other zipped files contain image data taken at PROMES-CNRS laboratory in Odeillo. Each zipped file contains all images for one day (the date is given in the file name).</li> </ul> <p>Images and GHI measures from the following days have been used for training and cross-validation:</p> <ol> <li>2020-09-11</li> <li>2020-09-12</li> <li>2020-09-16</li> <li>2020-09-19</li> <li>2020-09-20</li> <li>2020-09-21</li> <li>2020-09-22</li> <li>2020-09-23</li> <li>2020-09-24</li> <li>2020-09-29</li> <li>2020-10-01</li> </ol> <p>Images and GHI measures from the following days have been used for test:</p> <ol> <li>2020-10-04</li> <li>2020-10-05</li> <li>2020-10-08</li> <li>2020-11-05</li> <li>2020-11-15</li> </ol>
Data from: Skyline fossilized birth-death model is robust to violations of sampling assumptions in total-evidence dating
<p>Several total-evidence dating studies under the fossilized birth-death (FBD) model have produced very old age estimates, which are not supported by the fossil record. This phenomenon has been termed "deep root attraction (DRA)". For two specific datasets, involving divergence time estimation for the early radiations of ants, bees and wasps (Hymenoptera) and of placental mammals (Eutheria), it has been shown that the DRA effect can be greatly reduced by accommodating the fact that extant species in these trees have been sampled to maximize diversity, so called diversified sampling. Unfortunately, current methods to accommodate diversified sampling only consider the extreme case where it is possible to identify a cut-off time such that all splits occurring before this time are represented in the sampled tree but none of the younger splits. In reality, the sampling bias is rarely this extreme, and may be difficult to model properly. Similar modeling challenges apply to the sampling of the fossil record. This raises the question of whether it is possible to find dating methods that are more robust to sampling biases. Here, we show that the skyline FBD (SFBD) process, where the diversification and fossil-sampling rates can vary over time in a piecewise fashion, provides age estimates that are more robust to inadequacies in the modeling of the sampling process and less sensitive to DRA effects. In the SFBD model we consider, rates in different time intervals are either considered to be independent and identically distributed, or assumed to be autocorrelated following an Ornstein-Uhlenbeck (OU) process. Through simulations and reanalyses of the Hymenoptera and Eutheria data, we show that both variants of the SFBD model unify age estimates under random and diversified sampling assumptions. The SFBD model can resolve DRA by absorbing the deviations from the sampling assumptions into the inferred dynamics of the diversification process over time. Although this means that the inferred diversification dynamics must be interpreted with caution, taking sampling biases into account, we conclude that the SFBD model represents the most robust approach available currently for addressing DRA in total-evidence dating.</p>
A hierarchical graph-based model for mobility data representation and analysis
<p>Hierarchical representations of transportation networks should provide a better understanding of mobility patterns and the underlying structures at various abstraction levels. A hierarchical graph-based model allows representing moving objects and trajectories according to multiple spatial, temporal and semantic scales. The latter model is implemented here in a Neo4j graph database (version 4.4.0) and experimented with historical maritime data covering Brittany Bay in France.</p>
Data generated by the model presented in the research article entitled "Simulation of mass and heat transfer in an evaporatively cooled PEM fuel cell"
<p>This repository provides all the data and scripts necessary to reproduce the line plots shown in the manuscript entitled "Simulation of mass and heat transfer in an evaporatively cooled PEM fuel cell".</p>
Environmental data at the sampling event level collected with Inline instruments, almanach, models and satellites during the Tara Pacific Expedition 2016-2018
<p>The Tara Pacific expedition (2016-2018) sampled coral ecosystems around 32 islands in the Pacific Ocean, and sampled the surface of oceanic waters at 249 locations, resulting in the collection of nearly 58,000 samples. The expedition was designed to systematically study corals, fish, plankton, and seawater, and included the collection of samples for advanced biogeochemical, molecular, and imaging analysis. Here we provide at the sampling event level, the environmental data originating from all instruments acquiring continuously during the full course of the campaign. This dataset is augmented with the addition of variables originating from almanach (local sun/moon set/rise, local zenith), from operational models obtained from Copernicus Marine Services, but also <strong>f</strong>rom satellite imagery (MODIS-AQUA satellite - Level 3 mapped product, 8 day average, 4km resolution) at <a href="https://oceandata.sci.gsfc.nasa.gov">https://oceandata.sci.gsfc.nasa.gov</a>. The zone corresponding to the station position and date was recovered either by taking a two pixel buffer around the given location (total zone being a 5 by 5 pixels square of 20 km side) and in order to propose an alternative measure in the inevitable case where clouds were present an alternative 12 pixels buffer was taken (total zone being a 25 by 25 pixels square of 100 km side). All data were provided as mean, standard deviation (sd) together with 0.05, 0.25, 0.5, 0.75 and 0.95 quartiles</p>
Finite Element model data for Academic Rotor bladed-disc system
<div> <div> <div> <p>A computational finite element based technique is proposed for developing a stochastic reduced order model for rotating bladed disc with spatial random inhomogeneities. The spatial inhomogeneities imply the system to be randomly mistuned. The formulation assumes the availability of a high fidelity finite element (FE) model for the tuned system. The corresponding FE matrices are antisymmetric on account of the Coriolis forces due to rotation. The spatial inhomogeneities, available from limited point measurements on the blades, are modelled as non-Gaussian random fields with arbitrary distributions. A low order stochastic computational model is developed by projecting the FE model onto a reduced dimensional state space defined in terms of specified observable nodal points and expressing the stochasticity through an arbitrary polynomial chaos (aPC) basis. This model enables probabilistic quantification of the variabilities in the system response and estimating failure probabilities. The methodology enables drastic reduction in the state space and stochastic dimensions, addresses the practical difficulties with having limited measurable data points, antisymmetric FE matrices, aPC representation in complex irregular geometries and carrying out probabilistic analyses on industrial systems, at significantly reduced computational costs. The methodology is illustrated through an academic rotor and an industrial rotor blade.</p> </div> </div> </div>
Data of curvature model for the study of nanoparticle size effects on amyloid fibril stability and molecular dynamics simulations data
<p>The data provided refer to our published article:</p> <p>T. John, J. Adler, C. Elsner, J. Petzold, M. Krueger, L.L. Martin, D. Huster, H.J. Risselada, B. Abel, Mechanistic insights into the size-dependent effects of nanoparticles on inhibiting and accelerating amyloid fibril formation, J. Colloid Interface Sci. 622 (2022), 804–818. <a href="https://doi.org/10.1016/j.jcis.2022.04.134">https://doi.org/10.1016/j.jcis.2022.04.134</a></p> <p>This article is accompanied by a 'Data in Brief' article that explains in more detail the use of the curvature model and our molecular dynamics (MD) simulations:</p> <p>T. John, L.L. Martin, H.J. Risselada, B. Abel, Curvature model for nanoparticle size effects on peptide fibril stability and molecular dynamics simulation data, Data Brief 45 (2022), 108598. <a href="https://doi.org/10.1016/j.dib.2022.108598">https://doi.org/10.1016/j.dib.2022.108598</a></p>
Thetis Baltic Sea simulation: model and observation data sets
<p>Model and observation data sets used in article "Adjoint-based optimization of a regional water elevation model".</p>
Data for Microkinetic modeling of the transient CO2 methanation with DFT-based uncertainties in a Berty reactor
<p>Dataset and scripts for the manuscript "Microkinetic modeling of the transient CO2 methanation with DFT-based uncertainties in a Berty reactor", which has been submitted for review. The file contains all the raw data and the evaluation of the experiments. Additionally, all scripts for the microkinetic model are provided to perform transient simulations with all 5000 methanation mechanisms investigated in the manuscript.</p>
Conflict over the eukaryote root resides in strong outliers, mosaics and missing data sensitivity of site-specific (CAT) mixture models
Abstract Phylogenetic reconstruction using concatenated loci ("phylogenomics" or "supermatrix phylogeny") is a powerful tool for solving evolutionary splits that are poorly resolved in single gene/protein trees (SGTs). However, recent phylogenomic attempts to resolve the eukaryote root have yielded conflicting results, along with claims of various artefacts hidden in the data. We have investigated these conflicts using two new methods for assessing phylogenetic conflict. ConJak uses whole marker (gene or protein) jackknifing to assess deviation from a central mean for each individual sequence, while ConWin uses a sliding window to screen for incongruent protein fragments (mosaics). Both methods allow selective masking of individual sequences or sequence fragments in order to minimize missing data, an important consideration for resolving deep splits with limited data. Analyses focused on a set of 76 eukaryotic proteins of bacterial-ancestry previously used in various combinations to assess the branching order among the three major divisions of eukaryotes: Amorphea (mainly animals, fungi and Amoebozoa), Diaphoretickes (most other well-known eukaryotes and nearly all algae) and Excavata, represented here by Discoba (Jakobida, Heterolobosea, and Euglenozoa). ConJak analyses found strong outliers to be concentrated in under-sampled lineages, while ConWin analyses of Discoba, the most under-sampled of the major lineages, detected potentially incongruent fragments scattered throughout. Phylogenetic analyses of the full data using an LG-gamma model support a Discoba sister scenario (neozoan-excavate root), which rises to 99-100% bootstrap support with data masked according to either protocol. However, analyses with two site-specific (CAT) mixture models yielded widely inconsistent results and a striking sensitivity to missing data. The neozoan-excavate root places Amorphea and Diaphoretickes as more closely related to each other than either is to Discoba, a fundamental relationship that should remain unaffected by additional taxa.
Ice core and model data for Moseid et al. 2022
<p>These datasets are used in the publication "Using ice cores to evaluate CMIP6 aerosol concentrations over the historical era" with the authors Kine Onsum Moseid, Michael Schulz, Anja Eichler, Margit<br> Schwikowski, Joseph R. McConnell, Dirk Olivi ́e, Alison S. Criscitiello, Karl J. Kreutz, and Michel Legrand.</p> <p>The paper is currently in review when this data is published.</p> <p>The excel sheet dataset contains sulfate and black carbon records from 15 ice cores as presented in the paper. </p> <p>One zip file contain the part of the data from NorESM2-LM experiments as described in the paper. Another dataset will be published to compliment this dataset. </p>
Supporting data for "Modeling the albedo neutron decay source of radiation belt electrons and protons"
<p>Data sets are provided in support of the publication to appear in JGR-Space Physics. They include tabulated values of computed albedo neutron flux above the atmosphere, and of resulting radiation belt electron and proton source functions. Data format is described in the README files.</p>
Data from: The shape of aroma: Measuring and modeling citrus oil gland distribution
<p>From preventing scurvy to being part of religious rituals, citrus are intrinsically connected to human health and perception. From tiny mandarins to head-sized pummelos, citrus capability of hybridization provides a vastly diverse array of fruit sizes and shapes, which in turn corresponds to a diversity of flavors and aromas. These sensory qualities are tightly linked to oil glands in the citrus skin. The oil glands are also key to understanding fruit development, and the essential oils contained by them are fundamental in the food and perfume industries. We study the shape of citrus based on 3D X-ray CT scan reconstruction of 163 different citrus samples comprising 58 different species and cultivars, including samples of all fundamental citrus species. First, using the power of X-rays and image processing, we are able to compare and contrast size ratios between different tissues, such as the size of the skin compared to the rind or the flesh. Second, we model the fruit shape as an ellipsoidal surface, and later we study and infer possible oil gland distributions on this surface using principles of directional statistics. We finally compare and contrast these overall fruit shape models along their gland distributions across different citrus species. This morphological modeling will allow us later to link genotype with phenotype, furthering our insight on how the physical shape is genetically specified in DNA.</p>
Data from: A hierarchical model for jointly assessing ecological and anthropogenic impacts on animal demography
<p>1. The management of sustainable harvest of animal populations is of great ecological and conservation importance. Development of formal quantitative tools to estimate and mitigate the impacts of harvest on animal populations has positively impacted conservation efforts.</p> <p>2. The vast majority of existing harvest models, however, do not simultaneously estimate ecological and harvest impacts on demographic parameters and population trends. Given that the impacts of ecological drivers are often equal to or greater than the effects of harvest, and can covary with harvest, this disconnect has the potential to lead to flawed inference.</p> <p>3. In this study, we used Bayesian hierarchical models and a 43-year capture-mark-recovery dataset from 404,241 female mallards (Anas platyrhynchos) released in the North American midcontinent to estimate mallard demographic parameters. Further, we model the dynamics of waterfowl hunters and habitat, and the direct and indirect effects of anthropogenic and ecological processes on mallard demographic parameters.</p> <p>4. We demonstrate that density-dependence, habitat conditions, and harvest can simultaneously impact demographic parameters of female mallards, and discuss implications for existing and future harvest management models.</p> <p>5. Our results demonstrate the importance of controlling for multicollinearity among demographic drivers in harvest management models, and provide evidence for multiple mechanisms that lead to partial compensation of mallard harvest. We provide a novel model structure to assess these relationships that may allow for improved inference and prediction in future iterations of harvest management models across taxa.</p>
SDUST2020 MSS: A global 1′×1′ mean sea surface model determined from multi-satellite altimetry data
<p>SDUST2020 MSS (Shandong University of Science and Technology 2020 mean sea surface) model with a grid of 1′×1′ is established with 19-year moving average method from multi-satellite altimetry data over 27-year (from January 1993 to December 2019). Its spatial coverage is 80°S-84°N. The missions data of Topex/Poseidon, Jason-1, Jason-2, Jason-3, ERS-1, ERS-2, GFO, Envisat, SARAL, HY-2A, Sentinel-3A and Cryosat-2 are ingested in the SDUST2020 MSS model.</p>
Deep Learning for Reaction-Diffusion Glioma Growth Modeling: Towards a Fully Personalized Model? — Supporting Data
<p>Supporting data for Martens et al. Deep Learning for Reaction-Diffusion Glioma Growth Modelling: Towards a Fully Personalised Model? arXiv:2111.13404.</p>
Global demand data for PyPSA-Earth: An Open Optimisation Model of the Earth Energy System.
<p><strong>PyPSA-Earth </strong>is an open model dataset of the global power system at different network levels that cover our Earth. The African model can be built using the code provided at <a href="https://github.com/pypsa-meets-africa/pypsa-africa">https://github.com/pypsa-meets-africa/pypsa-africa</a>. Other regions follow soon under the same code base.</p> <p>Since the GitHub codebase is not suited for handling large changing files, we provide here separate <strong>data bundles and cutouts</strong> to be downloaded and extracted as noted in the <a href="https://pypsa-meets-africa.readthedocs.io/en/latest/index.html">documentation</a></p> <p>The below-provided <strong>resource file </strong>contains demand time-series generated by <a href="https://github.com/niclasmattsson/GlobalEnergyGIS/blob/b23206f8701acafdf7359f9cc952dfd4e7b819e5/src/downloaddatasets.jl">GEGIS</a> covering the world. The time series are produced for different socio-economic scenarios (SSP), weather years, and prediction years<strong>.</strong></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.