Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
66
datasets available to search
ShareScore release 0.9.0
Dataset results
66 results for “stochastic models”
Data from: A stochastic generative model for citation networks among academic papers
<p>We propose a stochastic generative model to represent a directed graph constructed by citations among academic papers, where nodes and directed edges represent papers with discrete publication time and citations respectively. The proposed model assumes that a citation between two papers occurs with a probability based on the type of the citing paper, the importance of cited paper, and the difference between their publication times, like the existing models. We consider the out-degrees of citing paper as its type, because, for example, survey paper cites many papers. We approximate the importance of a cited paper by its in-degrees. In our model, we adopt three functions: a logistic function for illustrating the numbers of papers published in discrete time, an inverse Gaussian probability distribution function to express the aging effect based on the difference between publication times, and an exponential distribution (or a generalized Pareto distribution) for describing the out-degree distribution. We consider that our model is a more reasonable and appropriate stochastic model than other existing models and can perform complete simulations without using original data. In this paper, we first use the Web of Science database and see the features used in our model. By using the proposed model, we can generate simulated graphs and demonstrate that they are similar to the original data concerning the in- and out-degree distributions, and node triangle participation. In addition, we analyze two other citation networks derived from physics papers in the arXiv database and verify the effectiveness of the model.</p>
Codes: An early warning indicator trained on stochastic disease-spreading models with different noises
<p>This dataset contains the training data (Version V1) and all the codes (Version V2) of the paper entitled "An early warning indicator trained on stochastic disease-spreading models with different noises." </p> <p>Time series and corresponding residuals from white noise (equation 2.5), environmental noise (equation 2.8), and demographic noise (equation 2.9) are stored in the training_data_WhiteN, training_data_EnvN, and training_data_DemN folders, respectively. All residuals of the time series are contained in the training_resids folder, which also includes labels and groups of the training data. For details on the data generation process, please refer to section 3.1 in the paper. </p>
Datasets for reproducing the results in "True random number generators with flicker noise: stochastic model, min-entropy calculation and online test"
<p>Datasets for reproducing the results in "True random number generators with flicker noise: stochastic model, min-entropy calculation and online test"</p> <p>Includes scripts for generating raw results, postprocessing scripts, as well as scripts/notebooks for generating plots for the manuscript</p>
Dataset for "Optimization and Evaluation of Stochastic Unified Convection Using Single-Column Model Simulations at Multiple Observation Sites"
<p>SCAM5 and LES outputs from "Optimization and Evaluation of Stochastic Unified Convection Using Single-Column Model Simulations at Multiple Observation Sites". Includes simulation outputs of stochastic UNICON and original UNICON. The LES intercomparison data of DYCOMSRF01 is available at https://gcss-dime.giss.nasa.gov/pub/DYCOMS-II/GCSS7-RF01/gcss7.nc, and the data of CGILS is available at http://www.atmos.washington.edu/~bloss/CGILS2data.tar.</p>
Dataset for paper Pavel Perezhogin, Laure Zanna, Carlos Fernandez-Granda "Generative data-driven approaches for stochastic subgrid parameterizations in an idealized ocean model" submitted to JAMES.
<p>The dataset consists of the directory tree of .zarr archives. See <a href="https://github.com/m2lines/pyqg_generative/blob/master/Google-Colab/dataset.ipynb">Github repository</a> for the description of the dataset.</p> <p>The directory tree is:</p> <pre><code>├── eddy │ ├── 48 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 64 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 96 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ └── hires ├── jet │ ├── 48 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 64 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 96 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ └── hires</code></pre> <ul> <li>Every individual dataset is a <code>.zarr</code> <a href="https://zarr.readthedocs.io/en/stable/">archive</a></li> <li><code>eddy/jet</code> - configuration of the pyqg; eddy is default; See <a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2022MS003258">Ross2022</a> for description</li> <li><code>hires.zarr</code> - high-resolution simulation at 256x256 grid</li> <li><code>48/64/96</code> - resolution of the coarse models</li> <li><code>lores.zarr</code> - low-resolution simulation</li> <li><code>gauss.zarr</code>, <code>sharp.zarr</code> - training datasets for prediction of subgrid forcing obtained with Gaussian or Sharp filters</li> <li><code>hires-gauss.zarr</code>, <code>hires-sharp.zarr</code> - high-resolution simulation projected onto coarse grid with Gaussian or Sharp filters</li> </ul> <p>The directory tree is split into small tar.gz files each representing a separate .zarr archive. Download any required parts of the dataset and unpack with:</p> <p><strong>tar -xf *.tar.gz </strong></p> <p><strong>The directory tree will be restored automatically!</strong></p>
Data from: A stochastic generative model for citation networks among academic papers
Open the record for dataset details and reuse information.
Stochastic Modeling of Subglacial Topography Exposes Uncertainty in Water Routing at Jakobshavn Glacier
<p>Abstract:</p> <p>These data products accompany the paper "Stochastic Modeling of Subglacial Topography Exposes Uncertainty in Water Routing at Jakobshavn Glacier" (MacKie et al., in review). In this study, geostatistical techniques were used to generate an ensemble of topographic realizations that retain the spatial statistics of radar bed elevation measurements. The simulation was conditioned to local radar data and mass conservation bed estimates. This repository contains the radar and mass conservation conditioning data, the ensemble of topographic realizations, and coordinate data.</p> <p> </p> <p>Content and processing steps:</p> <p>The study area is 75.15 x 48.90 km^2. The grid cell resolution is 150 meters. Each digital elevation model (DEM) has 501 x 326 grid cells. The mass conservation DEM was obtained from BedMachine Greenland (Morlighem and others, 2017). The radar data were acquired from the Center for Remote Sensing of Ice Sheets (CReSIS) 2009 flights (Gogineni, 2012; Gogineni and others, 2014). A probabilistic modeling technique called sequential Gaussian co-simulation (Verly, 1993; Almeida and Journel, 1994; Journel, 1999; Remy, 2005) was used to generate the topographic realizations. The datasets are as follows:</p> <p> </p> <p>1) Jakobshavn_mass_conservation.txt - Mass conservation conditioning data</p> <p>2) Jakobshavn_radar_data.txt - Radar conditioning data</p> <p>3) Jakobshavn_simulation.txt - 250 topographic realizations. The shape of this file is 250 x 163326, where each column corresponds to one topographic realization. Each column should be reshaped to 501 x 326 to view the DEM.</p> <p>4) Jakobshavn_x_data.txt - Polar stereographic X coordinates in meters</p> <p>5) Jakobshavn_y_data.txt - Polar stereographic Y coordinates in meters</p> <p> </p> <p>References:</p> <p>Almeida, A. S., & Journel, A. G. (1994). Joint simulation of multiple variables with a Markov-type coregionalization model. <em>Mathematical Geology</em>, <em>26</em>(5), 565-588.</p> <p>Gogineni, P. (2012). CReSIS radar depth sounder data. <em>Center for Remote Sensing of Ice Sheets, Lawrence, KS https://data. cresis.-ku. edu</em>.</p> <p>Gogineni, S., Yan, J. B., Paden, J., Leuschen, C., Li, J., Rodriguez-Morales, F., ... & Gauch, J. (2014). Bed topography of Jakobshavn Isbræ, Greenland, and Byrd Glacier, Antarctica. <em>Journal of Glaciology</em>, <em>60</em>(223), 813-833.</p> <p>Journel, A. G. (1999). Markov models for cross-covariances. <em>Mathematical Geology</em>, <em>31</em>(8), 955-964.</p> <p>Morlighem, M., Williams, C. N., Rignot, E., An, L., Arndt, J. E., Bamber, J. L., ... & Fenty, I. (2017). BedMachine v3: Complete bed topography and ocean bathymetry mapping of Greenland from multibeam echo sounding combined with mass conservation. <em>Geophysical research letters</em>, <em>44</em>(21), 11-051.</p> <p>Remy, N. (2005). S-GeMS: the stanford geostatistical modeling software: a tool for new algorithms development. In <em>Geostatistics banff 2004</em> (pp. 865-871). Springer, Dordrecht.</p> <p>Verly, G. W. (1993). Sequential Gaussian cosimulation: a simulation method integrating several types of information. In <em>Geostatistics Troia’92</em> (pp. 543-554). Springer, Dordrecht.</p>
Data from: A stochastic model for annual reproductive success
Demographic stochasticity can have large effects on the dynamics of small populations as well as on the persistence of rare genotypes and lineages. Survival is sensibly modeled as a binomial process, but annual reproductive success (ARS) is more complex and general models for demographic stochasticity do not exist. Here we introduce a stochastic model framework for ARS and illustrate some of its properties. We model a sequence of stochastic events: nest completion, the number of eggs or neonates produced, nest predation, and the survival of individual offspring to independence. We also allow multiple nesting attempts within a breeding season. Most of these components can be described by Bernoulli or binomial processes; the exception is the distribution of offspring number. Using clutch and litter size distributions from 53 vertebrate species, we demonstrate that among‐individual variability in offspring number can usually be described by the generalized Poisson distribution. Our model framework allows the demographic variance to be calculated from underlying biological processes and can easily be linked to models of environmental stochasticity or selection because of its parametric structure. In addition, it reveals that the distributions of ARS are often multimodal and skewed, with implications for extinction risk and evolution in small populations.
Data from: Experiments and modelling of rate-dependent transition delay in a stochastic subcritical bifurcation
Complex systems exhibiting critical transitions when one of their governing parameters varies are ubiquitous in nature and in engineering applications. Despite a vast literature focusing on this topic, there are few studies dealing with the effect of the rate of change of the bifurcation parameter on the tipping points. In this work, we consider a subcritical stochastic Hopf bifurcation under two scenarios: the bifurcation parameter is first changed in a quasi-steady manner and then, with a finite ramping rate. In the latter case, a rate-dependent bifurcation delay is observed and exemplified experimentally using a thermoacoustic instability in a combustion chamber. This delay increases with the rate of change. This leads to a state transition of larger amplitude compared to the one that would be experienced by the system with a quasi-steady change of the parameter. We also bring experimental evidence of a dynamic hysteresis caused by the bifurcation delay when the parameter is ramped back. A surrogate model is derived in order to predict the statistic of these delays and to scrutinise the underlying stochastic dynamics. Our study highlights the dramatic influence of a finite rate of change of bifurcation parameters upon tipping points and it pinpoints the crucial need of considering this effect when investigating critical transitions.
Data from: Hypothesised diprotomeric enzyme complex supported by stochastic modelling of Palytoxin-induced Na/K pump channels
The sodium-potassium pump (Na+/K+ pump) is crucial for cell physiology. Despite great advances in the understanding of this ionic pumping system, its mechanism is not completely understood. We propose the use of the Statistical Model Checker to investigate palytoxin-induced Na+/K+ pump channels. We modelled a system of reactions representing transitions between the conformational substates of the channel with parameters, concentrations of the substates, and reaction rates extracted from simulations reported in the literature, based on electrophysiological recordings in a whole-cell configuration. The model was implemented using the UPPAAL-SMC platform. Comparing simulations and probabilistic queries from stochastic system semantics with experimental data, it was possible to propose additional reactions to reproduce the single channel dynamic. The probabilistic analyses and simulations suggest that the palytoxin-induced Na+/K+ pump channel functions as a diprotomeric complex in which protein-protein interactions increase the affinity of the Na+/K+ pump affinity for palytoxin.
Data from: Accounting for uncertainty in dormant life stages in stochastic demographic models
Dormant life stages are often critical for population viability in stochastic environments, but accurate field data characterizing them are difficult to collect. Such limitations may translate into uncertainties in demographic parameters describing these stages, which then may propagate errors in the examination of population-level responses to environmental variation. Expanding on current methods, we 1) apply data-driven approaches to estimate parameter uncertainty in vital rates of dormant life stages and 2) test whether such estimates provide more robust inferences about population dynamics. We built integral projection models (IPMs) for a fire-adapted, carnivorous plant species using a Bayesian framework to estimate uncertainty in parameters of three vital rates of dormant seeds – seed-bank ingression, stasis and egression. We used stochastic population projections and elasticity analyses to quantify the relative sensitivity of the stochastic population growth rate (log λs) to changes in these vital rates at different fire return intervals. We then ran stochastic projections of log λs for 1000 posterior samples of the three seed-bank vital rates and assessed how strongly their parameter uncertainty propagated into uncertainty in estimates of log λs and the probability of quasi-extinction, Pq(t). Elasticity analyses indicated that changes in seed-bank stasis and egression had large effects on log λs across fire return intervals. In turn, uncertainty in the estimates of these two vital rates explained > 50% of the variation in log λs estimates at several fire-return intervals. Inferences about population viability became less certain as the time between fires widened, with estimates of Pq(t) potentially > 20% higher when considering parameter uncertainty. Our results suggest that, for species with dormant stages, where data is often limited, failing to account for parameter uncertainty in population models may result in incorrect interpretations of population viability.
Data from: Pathogen growth in insect hosts: inferring the importance of different mechanisms using stochastic models and response time data
Open the record for dataset details and reuse information.
Data from: A stochastic model for annual reproductive success
Open the record for dataset details and reuse information.
Data from: A stochastic neuronal model predicts random search behaviors at multiple spatial scales in C. elegans
Open the record for dataset details and reuse information.
Data from: Experiments and modelling of rate-dependent transition delay in a stochastic subcritical bifurcation
Open the record for dataset details and reuse information.
Supplementary code for: Polygenic local adaptation in metapopulations: a stochastic eco-evolutionary model
Open the record for dataset details and reuse information.
Data from: A stochastic vision based model inspired by the collective behaviour of zebrafish in heterogeneous environments
Open the record for dataset details and reuse information.
Data from: Hypothesised diprotomeric enzyme complex supported by stochastic modelling of Palytoxin-induced Na/K pump channels
Open the record for dataset details and reuse information.
Data from: Accounting for uncertainty in dormant life stages in stochastic demographic models
Open the record for dataset details and reuse information.
Stochastic Persistence in Nucleated Polymerization Model
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.