Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Data for "Measurement report: Comparison of airborne in-situ measured, lidar-based, and modeled aerosol optical properties in the Central European background – identifying sources of deviations"
<p>A unique set of data is presented, derived from measurements conducted at the rural central European observatory at Melpitz, Germany. Data derived from remote sensing (lidar), airborne platforms (helicopter, balloon), and ground-based in-situ methods is included. Measured and Mie-modeled optical aerosol parameters are presented in the dry- and ambient state. Modeled optical parameters are based on Mie-theory. For ambient state hygroscopic growth simulations are utilized.</p>
ROM model data for North Western Mediterranean Deep Water Formation
<p>ROM model data for North Western Mediterranean Deep Water Formation</p>
Data from: N-mixture models estimate abundance reliably: a field test on Marsh Tit using time-for-space substitution
<p>Imperfect detection in field studies on animal abundance, including birds, is common and can be corrected for in various ways. The binomial N-mixture (hereafter binmix) model developed for this task is widely used in ecological studies owing to its simplicity: it requires replicated count results as the input. However, it may overestimate abundance and be sensitive to even small violations of its assumptions. We used a 33-year dataset on the Marsh Tit, Poecile palustris, a sedentary forest passerine, from Białowieża Forest, Poland to validate inference from binmix models by comparing model-estimated abundances to the true number of breeding pairs within the plots, determined by exhaustive population study. The abundance estimates, derived from six springtime (April-May) counts of males on each plot in each year, were highly reliable: 116 out of 132 year-plot estimates (88%) included the true number of pairs within the 95% confidence intervals. Over- and underestimations were thus rare and similarly frequent (9 and 12 cases, respectively), with a tendency to overestimate at low densities and underestimate at high densities. Marsh Tits sing rarely but the frequency of countersinging increases with abundance, leading to non-independence in detections. When accounted for in a submodel for detection, the per-survey number of countersinging events positively affected detection probability but only weakly affected abundance estimates. Simulations further demonstrate that this property, overestimation at low densities and underestimation at high densities, may be a systematic bias of binmix model even if density-dependent detection is absent. While the behaviour of binmix models in specific situations requires more study, we conclude that these models are a valid tool to estimate abundance reliably when intensive population monitoring is not feasible.</p>
LNT Model is not an "Assumption": Re-Analysis of Epidemiological Data Empirically Supports LNT
<p><em>Introduction:</em> In “Keeping ICRP Recommendations Fit for Purpose”[1], LNT model is described as “LNT is the most appropriate evidence-based assumption to use for radiological protection purposes (p.10).” According to our critical literature survey on radiological epidemiology[2], some limitations were identified: (1) aggregation of individual level data, (2) model formulation, (3) model estimation, (4) model selection, (5) results interpretation. In this paper we focus (4) model selection and demonstrate LNT was the best model.</p> <p><em>Data and Method:</em> Using “Life Span Study Report 14. Cancer and non-cancer disease mortality data, 1950-2003 [3]”, solid cancer mortality was re-analyzed. In addition to the L, Q, LQ, hadn searched threshold model, kinked–at-2Gy model that assumes LQ for less than 2Gy and L for larger than 2 Gy, and Linear model with threshold as a parameter were estimated. Following [3], Poisson regression model was applied and model fit was compared with AIC and BIC.</p> <p><em>Results</em>: Among estimated models, Linear (BIC=18317.9) and grid search threshold at 20mSv (BIC=18318.1) was selected as the best models. Directly estimated threshold was -23.2 mSv and it was statistically insignificant (z=-0.087,p>0.1). Model fit of kinked-at-2Gy was poorer than these models (BIC=18321.2).</p> <p><em>Conclusion:</em> Based on these results, we can conclude LNT model is the best model for a-bomb survivor solid cancer mortality. According to our literature survey, LNT is supported Description of LNT model in “Keeping ICRP Recommendations Fit for Purpose” should be modified accordingly: “LNT is the scientifically supported model, it is reasonable LNT to use for radiological protection purposes.”</p>
Data and code to replicate: Diet analysis using generalized linear models derived from foraging processes using R package mvtweedie
<p>Diet analysis integrates a wide variety of visual, chemical and biological identification of prey. Samples are often treated as compositional data, where each prey is analyzed as a continuous percentage of the total. However, analyzing compositional data results in analytical challenges, e.g., highly parameterized models or prior transformation of data. Here, we present a novel approximation involving a Tweedie generalized linear model (GLM). We first review how this approximation emerges from considering predator foraging as a thinned and marked point process (with marks representing prey species and individual prey size). This derivation can motivate future theoretical and applied developments. We then provide a practical tutorial for the Tweedie GLM using new package <i>mvtweedie</i> that extends capabilities of widely used packages in R (<i>mgcv</i> and <i>ggplot2</i>) by transforming output to calculate prey compositions. We demonstrate this approach and software using two examples. Tufted puffins (<i>Fratercula cirrhata</i>) provisioning their chicks on a colony in the northern Gulf of Alaska show decadal prey switching among sand lance and prowfish (1980-2000) and then Pacific herring and capelin (2000-2020), while wolves (<i>Canis lupus ligoni</i>) in Southeast Alaska forage on mountain goats and marmots in northern uplands and marine mammals in seaward island coastlines. </p>
Study Data: Obtaining Semi-Formal Models from Qualitative Data: From Interviews into BPMN Models in User-Centered Design Processes
<p>This dataset (Data.zip) contains the raw data of a user study on the investigation of transforming think aloud interviews into BPMN models. All information on how to use the data are provide in the SPSS files and as a readme file. This transformation is executed following a manual additionally provided in Documents.zip. For the training phase, a website was used provided in Website.zip including Screenshots for simpler re-use. Further information are also included as readme file in the zip container.</p> <p>Main research question answered is in how far the manual reduces interpretation and variance in the created models. </p>
Data from: Human face-off: a new method for mapping evolutionary rates on three-dimensional digital models
<p>Modern phylogenetic comparative methods allow estimating evolutionary rates of phenotypic change, how these rates differ across clades, and assessing whether the rate remained constant over time. Unfortunately, currently available phylogenetic comparative tools express the rate in terms of a scalar dimension, hence they do not allow us to determine rate variations among different parts of a single, complex phenotype, or charting of realized rate variation directly onto the phenotype. Herein, we present a new method which allows the mapping of evolutionary rate variation directly on three-dimensional phenotypes, informing on the direction and magnitude of trait change automatically.</p> <p>This new method, implemented by the function rate.map embedded in the R package 'RRphylo', is based on phylogenetic ridge regression rate estimates. Since the latter represent ridge regression slopes, they possess sign and magnitude. In 'RRphylo', different rates are calculated for different districts of the phenotype, which can then be visualized directly onto the phenotype itself. We present the application of rate.map to the evolution of facial skeleton in Hominoidea (the clade including living and fossil apes), the primate clade inclusive of Homo and the greater apes. We found that the highly derived, unique shape of the face in modern humans evolved through rapid phenotypic changes affecting the nasal bones, the brow ridge and the maxillary region. The canine fossa, a facial feature unique to Homo sapiens, did not belong to a region of rapid phenotypic change, and could be seen as the by-product of midface evolution as suggested by previous studies.</p>
Data for "Wave dispersion and dissipation in landfast ice: comparison of observations against models"
<p>Data to replicate Figures 2, 3 and 7 in the manuscript: Wave dispersion and dissipation in landfast ice: comparison of observations against models. Article submitted for review to The Cryosphere (https://doi.org/10.5194/tc-2021-210)</p>
Data from: Sequential use of niche and occupancy models identifies conservation and research priority areas for two data-poor endemic birds from the Colombian Andes
<p>The lack of high-quality information on data-poor species can hinder efforts to inform conservation actions via spatial distribution modeling. This is particularly true for tropical birds of conservation concern, for which ecological studies and assessments of their conservation status have received limited funding. Here we use a cost- and time-efficient protocol for assessing the distribution of range-restricted taxa and to identify priority areas for their conservation based on a sequential application of Environmental Niche Models (ENMs) and Occupancy-Detection Models. This approach first uses available geographical information and niche-theory to prioritize potential study sites, which can later be surveyed to obtain high-quality presence-absence data to accurately model distributional ranges with limited resources. We apply this protocol to identify priority areas for two Neotropical birds of conservation concern endemic to the Colombian Andes: Yellow-headed Brush-finch (<i>Atlapetes flaviceps</i>) and Tolima Dove (<i>Leptotila conoveri</i>). We first fitted ENMs using spatially-filtered datasets containing all available records up to 2018. We then conducted field surveys across climatically suitable areas identified for both species, carrying out a total of 1750 counts to generate input data for the occupancy models. Overall, our results suggested more extended and more continuous distribution ranges for both species than previously reported, but also identified population strongholds that are not currently represented within the national protected areas system. Both species occupied a narrow elevational belt (~1300–2600) of the Central Andes of Colombia primarily on the slopes of the Magdalena River valley, with isolated populations in the Western and Eastern Andes; these areas have undergone some of the most marked landscape transformations in Colombia. This straightforward protocol maximizes available information and minimizes costs, while allowing for estimation of occurrence probabilities for range-restricted, data-poor taxa.</p>
Data supplement for "Multiscale perspective on wetting on switchable substrates: mapping between microscopic and mesoscopic models"
<p>Data set and python code to recreate the figures of "Multiscale perspective on wetting on switchable substrates: mapping between microscopic and mesoscopic models". Additionally, it includes the oomph-lib Code to reproduce the data for the simulations in the mesoscopic thin-film model.</p>
Data used in the publication: Sensitivity of modeled microphysics to stochastically perturbed parameters
<p>These data support the results presented in the manuscript titled "Sensitivity of modeled microphysics to stochastically perturbed parameters". They consist of results from an idealized single vertical column atmospheric model run for a number of experiments that explore methods of representing model uncertainty. </p>
Data and scripts accompanying the paper "University of Warsaw Lagrangian Cloud Model (UWLCM) 2.0"
<p>The archive contains datasets, run scripts and plotting scripts used when preparing the paper:<br> P. Dziekan and P. Zmijewski "University of Warsaw Lagrangian Cloud Model (UWLCM) 2.0: Adaptation of a mixed Eulerian-Lagrangian numerical model for heterogeneous computing clusters"<br> submitted to Geoscientific Model Development on 19.11.2021.<br> </p>
Models of evolutionary rescue with plasticity and nonlinear environmental change: code and data
<p>Rapid environmental changes are putting numerous species at risk of extinction. For migration-limited species, persistence depends on either phenotypic plasticity or evolutionary adaptation (evolutionary rescue). Current theory on evolutionary rescue typically assumes linear environmental change. Yet accelerating environmental change may pose a bigger threat. Here we present the simulation code and data from a model of a species encountering an environment with accelerating or decelerating change, to which it can adapt through evolution or phenotypic plasticity (within-generational or transgenerational). We show that unless either form of plasticity is sufficiently strong or adaptive genetic variation is sufficiently plentiful, accelerating or decelerating environmental change increases extinction risk compared to linear environmental change for the same mean rate of environmental change. </p>
Data assimilation products by using multiple climate model simulations and different combinations of proxies
<p>This dataset of the climate reconstruction by data assimilation using isotope ratios provides annual surface air temperature, precipitation amount, and other climate variables during 850–2000.</p> <p>Two isotopes-incorporated atmospheric general circulation models and 129 isotopic proxy data (65 corals, 43 ice cores, and 21 tree-ring cellulose) were used in this study. There are nine experiments using three type of simulations and three combinations of proxies.</p> <p>The associated publication: Shoji, S., Okazaki, A., & Yoshimura, K. (2020). Impact of proxies and prior estimates on data assimilation using isotope ratios for the climate reconstruction of the last millennium. (submitted to Earth and Space Science)</p> <p>[Data structure]<br> X(lon) x Y(lat) x Z(2) x Variables(8) x Year(1151)<br> Z(1): analyses<br> Z(2): priors</p>
Data for Minimum air temperature modeling using RS data
<p>Data for Minimum air temperature modeling using RS data</p>
Monthly mean optical depth at 550 nm derived from AERONET data used for model evaluation in GMD-2021-357
<p>Climatological monthly means over 2000-2014 derived from Aerosol Robotic Network version 3 level 2.0 direct sun retrievals (monthly data) used for model evaluation in Myriokefalitakis et al. (2021), doi: 10.5194/gmd-2021-357</p> <p>The netCDF file includes the AOD at the 4 native AERONET wavelengths (440 nm, 670 nm, 870 nm and 1020 nm), as well as the interpolated values at 550 nm used for the evaluation. The statistics stored are calculated over the monthly values for the 15 year period and include: monthly mean, 5th, 50th, and 95th percentile, standard deviation, standard error, number of days available per station and month and number of days where coarse AOD dominates (used as proxy for dusty days). </p> <p>The AERONET retrievals of optical depth were downloaded through the AERONET data download tool (available at: https://aeronet.gsfc.nasa.gov/, last accessed March 28, 2020). We thank the principal investigators and their collaborators for their effort in establishing and maintaining all AERONET sites used in this compilation.</p> <p>The use of this dataset must follow the guidelines of the original data providers at AERONET, explained here: https://aeronet.gsfc.nasa.gov/new_web/data_usage.html</p>
Representing surface heterogeneity in land-atmosphere coupling in E3SMv1 single-column model over ARM SGP during summertime - E3SM SCM data and code
<p>This dataset contains post-processed E3SM single-column model output and code used to produce the figures in the manuscript that we are targeting Geoscientific Model Development to submit. </p>
Data to support publication figures and animation scripts at GitHub: Modeling weather-driven long-distance dispersal of spruce budworm moths (Choristoneura fumiferana)
<p>Long-term studies of insect populations in the North American boreal forest have shown the vital importance of long-distance dispersal to the maintenance and expansion of insect outbreaks. In this work, we extend several concepts established previously in an empirically-based dispersal flight model with recent work on the physiology and behavior of the adult eastern spruce budworm (SBW) moth, Choristoneura fumiferana (Clem.). An outbreak of defoliating SBW in Quebec, ongoing since the mid-2000s, already covers millions of hectares of forests in eastern Canada and threatens to spread into neighboring areas through annual summertime episodes of long-distance dispersal. Such flight events in favorable conditions frequently include billions of SBW moths dispersing in the warm atmospheric boundary layer, typically starting around sunset and often lasting through several hours of wind-driven transport over hundreds of kilometers. Successful SBW dispersal to possibly distant host forest areas depends acutely on the weather. Here we describe the components and results of SBW–pyATM, an open-source individual-based modeling framework developed in Python for the simulation of these weather-driven SBW dispersal events. Using seasonal SBW phenology results from BioSIM at known outbreak locations and high-resolution Weather Research and Forecasting (WRF) model output, we focus on modeling dispersal flights over two successive nights in July 2013 in southern Quebec. Our flight model closely reproduces the SBW spatial patterns and motions observed by weather surveillance radar over the St. Lawrence estuary. With SBW–pyATM we can estimate landing locations for both male and female SBW and the resulting spatial patterns of egg distribution, allowing us eventually to forecast future larval defoliation activity in new locations where immigration could help overcome local limitations on SBW populations. This information could then support forest management decisions where SBW outbreaks threaten valuable resources.</p>
Model and data for: Economical defense of resources structures territorial space use in a cooperative carnivore
<p>Manuscript Abstract: Ecologists have long sought to understand space use and mechanisms underlying patterns observed in nature. We developed an optimality landscape and mechanistic territory model to understand mechanisms driving space use and compared model predictions to empirical reality. We demonstrate our approach using gray wolves (<i>Canis lupus</i>). In the model, simulated animals selected territories to economically acquire resources by selecting patches with greatest value, accounting for benefits, costs, and tradeoffs of defending and using space on the optimality landscape. Our approach successfully predicted and explained first- and second-order space use of wolves, including the population's distribution, territories of individual packs, and influences of prey density, competitor density, human-caused mortality risk, and seasonality. It accomplished this using simple behavioral rules and limited data to inform the optimality landscape. Results contribute evidence that economical territory selection is a mechanistic bridge between space use and animal distribution on the landscape. This approach and resulting gains in knowledge enable predicting effects of a wide range of environmental conditions, contributing to both basic ecological understanding of natural systems and conservation. We expect this approach will demonstrate applicability across diverse habitats and species, and that its foundation can help continue to advance understanding of spatial behavior.</p> <p>Model & Data Abstract: In support of the above manuscript, all model files and data to re-create the analyses for the manuscript are included on Dryad. The model can be run in NetLogo (installation file included), using the associated input files to build the Montana landscape for wolves. Expertise in NetLogo is strongly recommended for using this model. Output files are likewise included along with code to create each plot in the manuscript and SI. Software files for the model and code to create each plot in the manuscript are located at Zenodo: https://doi.org/10.5281/zenodo.5802243.</p>
Data and code for: Fixed effects or random effects in statistical models? fewer than five levels of a grouping factor
<p>This is code (and simulated data from that code) to assess how sample size and the numbers of levels of random effects influence parameter estimates of fixed effects in linear mixed-effects models. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.