Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
144
datasets available to search
ShareScore release 0.9.0
Dataset results
144 results for “statistical modeling”
Data associated with ecological niche models and post-ENM statistical analyses for Trillium species distributions
Open the record for dataset details and reuse information.
Data from: High quality statistical shape modelling of the human nasal cavity and applications
Open the record for dataset details and reuse information.
Data from: Maiasaura, a model organism for extinct vertebrate population biology: a large sample statistical assessment of growth dynamics and survivorship
Open the record for dataset details and reuse information.
Data from: Statistical comparison of trait-dependent biogeographical models indicates that Podocarpaceae dispersal is influenced by both seed cone traits and geographical distance
Open the record for dataset details and reuse information.
Data from: A statistical mechanics framework for constructing non-equilibrium thermodynamic models
Open the record for dataset details and reuse information.
Data from: Use of simulation-based statistical models to complement bioclimatic models in predicting continental scale invasion risks
Open the record for dataset details and reuse information.
Data from: A statistical skull geometry model for children 0-3 years old
Open the record for dataset details and reuse information.
Data from: The Cumulative Indel Model: fast and accurate statistical evolutionary alignment
Open the record for dataset details and reuse information.
Calibration of probability predictions from machine-learning and statistical models
Open the record for dataset details and reuse information.
SuperDARN grid files used to create the Thomas and Shepherd [2018] statistical convection model
<p>Daily Northern Hemisphere SuperDARN grid files in DataMap format for the years 2010-2016 (inclusive) used to create the <em>Thomas and Shepherd</em> [2018] statistical convection model. Grid files were processed using a modified copy of v4.2 of the Radar Software Toolkit (RST) where the average slant range associated with each grid vector is stored in the 'pwr.median' field.</p> <p>References:</p> <ul> <li>Thomas, E. G., and S. G. Shepherd (2018), Statistical patterns of ionospheric convection derived from mid-latitude, high-latitude, and polar SuperDARN HF radar observations, J. Geophys. Res. Space Physics, 123, 3196-3216, doi:10.1002/2018JA025280.</li> <li>SuperDARN Data Analysis Working Group. Participating members: Thomas, E. G., P. V. Ponomarenko, D. D. Billett, E. C. Bland, A. G. Burrell, K. Kotyk, A. S. Reimer, M. Schmidt, S. G. Shepherd, K. T. Sterne, and M.-T. Walach, (2018), SuperDARN Radar Software Tookit (RST) v4.2, Zenodo, doi:10.5281/zenodo.1403226.</li> </ul>
Model Outputs for Lobster Farm location Statistics
<p>Results of the Delft3D model for summer and winter scenarios to look at how the current velocities vary between the different areas of the lobster farm, results compared using ANOVA and post hoc statistical tests.</p>
Data from: Learning to count: determining the stoichiometry of bio-molecular complexes using fluorescence microscopy and statistical modelling
<p>As stated in the Read Me file:</p> <p>These data and resources are associated with the manuscript:</p> <p><em>Learning to count: determining the stoichiometry of bio-molecular complexes using fluorescence microscopy and statistical modelling</em>, Mersmann et. al., as submitted to biorXiv in July 2020.</p> <p>The raw imaging data relates to Figure 5, S1, S2 and Table S1. The images are fluorescent micrographs displaying immobilised adenovirus particles bound to a monoclonal antibody 9C12.</p> <p>Each experiment folder is numbered, as in Table S1, and appended with the mixing proportion (Fl), as defined in the manuscript. Within each folder there are 6 subfolders, representing samples incubated with different concentrations of 9C12 antibody.</p> <p>Each image is a 3 channel 1024x1024 tif. Channel 1 = 9C12 Alexa Fluor 647. Channel 2 = 9C12 Biotin + QDot655. Channel 3 = Adenovirus Alexa Fluor 488. Samples were illuminated in TIRF mode using a 100X objective, images were captured on a Hamamatsu OCRA Flash 4 sCMOS camera. Further details are available in the header of each file.</p> <p>The control samples are labelled with 100% 9C12 Alexa Fluor 647 or 100% 9C12 Biotin, as described in the manuscript.</p> <p>The data analysis script is an imageJ macro. It runs on the FIJI version of ImageJ with the NanoJ package installed (https://github.com/HenriquesLab). It outputs fluorescent measurements for each identified AdV particle. Note that the script rearranges the channel order such that Channel 1 = Adenovirus Alexa Fluor 488, Channel 2 = 9C12 Alexa Fluor 647, Channel 3 = 9C12 Biotin + QDot655. </p> <p>The channels require registration due to chromatic aberration, this is achieved using the Realign Channels function in NanoJ, appropriate translation masks are provided along with the script.</p> <p>Any question about the data or script should be addressed in Joe Grove (j.grove@ucl.ac.uk)</p>
Bayesian inference of tree species using diffusion models: tabulated posterior statistics for SNAPP and SNAPPER analyses
<p>We describe a new and computationally efficient Bayesian methodology for inferring species trees and demographics from unlinked binary markers. Likelihood calculations are carried out using diffusion models of allele frequency dynamics combined with novel numerical algorithms. The diffusion approach allows for analysis of datasets containing hundreds or thousands of individuals. The method, which we call \snapper, has been implemented as part of the BEAST2 package. We conducted simulation experiments to assess numerical error, computational requirements and accuracy recovering known model parameters. A re-analysis of soybean SNP data demonstrates that the models implemented in \snapp and \snapper can be difficult to distinguish in practice, a characteristic which we tested with further simulations. We demonstrate the scale of analysis possible using a SNP dataset sampled from 399 fresh water turtles in 41 populations.</p>
Data from: Evaluating predictive performance of statistical models explaining wild bee abundance in a mass-flowering crop
<p>Wild bee populations are threatened by current agricultural practices in many parts of the world, which may put pollination services and crop yields at risk. Loss of pollination services can potentially be predicted by models that link bee abundances with landscape-scale land-use, but there is little knowledge on the degree to which these statistical models are transferable across time and space. This study assesses the transferability of models for wild bee abundance in a mass-flowering crop across space (from one region to another) and across time (from one year to another). The models used existing data on bumblebee and solitary bee abundance in winter oilseed rape fields, together with high-resolution land-use crop-cover and semi-natural habitats data, from studies conducted in five different regions located in four countries (Sweden, Germany, Netherlands, and the UK), in three different years (2011, 2012, 2013). We developed a hierarchical model combining all studies and evaluated the transferability using cross-validation. We found that both the landscape-scale cover of mass-flowering crops and permanent semi-natural habitats, including grasslands and forests, are important drivers of wild bee abundance in all regions. However, while the negative effect of increasing mass-flowering crops on the density of the pollinators is consistent between studies, the direction of the effect of semi-natural habitat is variable between studies. The transferability of these statistical models is limited, especially across regions, but also across time. Our study demonstrates the limits of using statistical models in conjunction with widely available land-use crop-cover classes for extrapolating pollinator density across years and regions, likely in part because input variables such as cover of semi-natural habitats poorly capture variability in pollinator resources between regions and years.</p>
Data from: Modelling competition and dispersal in a statistical phylogeographic framework
Competition between organisms influences the processes governing the colonization of new habitats. As a consequence, species or populations arriving first at a suitable location may prevent secondary colonization. While adaptation to environmental variables (e.g., temperature, altitude, etc.) is essential, the presence or absence of certain species at a particular location often depends on whether or not competing species co-occur. For example, competition is thought to play an important role in structuring mammalian communities assembly. It can also explain spatial patterns of low genetic diversity following rapid colonization events or the "progression rule" displayed by phylogenies of species found on archipelagos. Despite the potential of competition to maintain populations in isolation, past quantitative analyses have largely ignored it because of the difficulty in designing adequate methods for assessing its impact. We present here a new model that integrates competition and dispersal into a Bayesian phylogeographic framework. Extensive simulations and analysis of real data show that our approach clearly outperforms the traditional Mantel test for detecting correlation between genetic and geographic distances. But most importantly, we demonstrate that competition can be detected with high sensitivity and specificity from the phylogenetic analysis of genetic variation in space.
Data from: Spatially structured statistical network models for landscape genetics
A basic understanding of how the landscape impedes, or creates resistance to, the dispersal of organisms and hence gene flow is paramount for successful conservation science and management. Spatially structured ecological networks are often used to represent spatial landscape-genetic relationships, where nodes represent individuals or populations and resistance to movement is represented using non-binary edge weights. Weights are typically assigned or estimated by the user, rather than observed, and validating such weights is challenging. We provide a synthesis of current methods used to estimate edge weights and an overview of common model types, stressing the advantages and disadvantages of each approach and their ability to model landscape-genetic data. We further explore a set of spatial-statistical methods that provide ecologists with alternative approaches for modeling spatially explicit processes that may affect genetic structure. This includes an overview of spatial autoregressive models, with a particular focus on how correlation and partial correlation are used to represent neighborhood structure with the inverse of the covariance matrix (i.e., precision matrix). We then demonstrate how to model resistance by specifying an appropriate statistical model on the nodes, conditioned on the edge weights, through the precision matrix. This integration of network ecology and spatial statistics provides a practical analytical framework for landscape-genetic studies. The results can be used to make statistical inferences about the relative importance of individual landscape characteristics, such as the vegetative cover, hillslope, or the presence of roads or rivers, on gene flow. In addition, the R code we include allows readers to explore landscape-genetic structure in their own datasets, which will potentially provide new insights into the evolutionary processes that generated ecological networks, as well as valuable information about the optimal characteristics of conservation corridors.
Data from: The relationship between native species richness and exotic species richness or occurrence will always be negative when the total number of species is accounted for in statistical models: A response to Beaury et al.
Beaury et al. (2020) attempt to address the scale dependence of evidence for biotic resistance by including environmental covariates that can account for total species richness. However, this approach will incorrectly estimate relationships, driven by the accuracy of the covariates rather than the true relationship between native and non-native species.
Data for Herb-paths, a network and statistical model to explore health-beneficial effects of herbs and herbal constituents
<p>Results data for the manuscript "Herb-paths, a network and statistical model to explore health-beneficial effects of herbs and herbal constituents".</p>
Forecasting Cryptocurrency Markets: Predictive Modelling Using Statistical and Machine Learning Approaches
Open the record for dataset details and reuse information.
ECONOMETRIC MODEL OF PUBLIC UTILITIES BASED ON MODERNIZATION AND UNIFICATION OF THE ECONOMY-STATISTICAL ANALYSIS
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.