Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

144

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

144 results for “statistical modeling”

Learn how ShareScore rates datasets ↗
dryad36/100

Raw data and R code for statistical analyses from: Sensory trap leads to reliable communication without a shift in nonsexual responses to the model cue

Open the record for dataset details and reuse information.

publicFeb 2024View details →
edi36/100

Settlement-Era Temperature Statistical Model, Upper Midwest: Level 2

Scientific records of temperature and precipitation have been kept for several hundred years, but for many areas, only a shorter record exists. To understand climate change, there is a need for rigorous statistical reconstructions of paleoclimate using proxy data. Paleoclimate proxy data are often sparse, noisy, indirect measurements of the climate process of interest, making each proxy uniquely challenging to model statistically. We reconstruct spatially-explicit temperature surfaces from sparse and noisy measurements recorded at historical United States military forts and other observer stations from 1820-1894. One common method for reconstructing paleoclimate from proxy data is principal component regression (PCR). With PCR, one learns a statistical relationship between the paleoclimate proxy data and a set of climate observations that are used as patterns for potential reconstruction scenarios. We explore PCR in a Bayesian hierarchical framework, extending classical PCR in a variety of ways. First, we model the latent principal components probabilistically, accounting for measurement error in the observational data. Next, we extend our method to better accommodate outliers that occur in the proxy data. Finally, we explore alternatives to the truncation of lower order principal components using different regularization techniques. One fundamental challenge in paleoclimate reconstruction efforts is the lack of out-of-sample data for predictive validation. Cross-validation is of potential value, but is computationally expensive and potentially sensitive to outliers in sparse data scenarios. To overcome the limitations that a lack of out-of-sample records presents, we test our methods using a simulation study, applying proper scoring rules including a computationally efficient approximation to leave-one-out cross-validation using the log score to validate model performance. The result of our analysis is a spatially explicit reconstruction of spatio-temporal temperat

openCC (other)Jan 2020View details →
zenodo32/100

Statistical model training data for "Continuous Structural Parameterization: A proposed method for representing different model parameterizations within one structure demonstrated for atmospheric convection"

<p>Gzipped CSV files containing convection scheme inputs and outputs used for training.</p> <p>Column format of each file:</p> <p>THETA_IN_1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,Q_IN_1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,DTHETA_1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,DQ_1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28</p> <p>where THETA_IN are input values of potential temperature [K], Q_IN are input values of specific humidity [kg/kg], DTHETA are changes in potential temperature due to convection [K], DQ are changes in specific humidity due to convection [kg/kg].</p> <p>Key:</p> <p>&quot;llcs&quot; are simulations with Lambert-Lewis.</p> <p>&quot;gr&quot; are simulations with Gregory-Rowntree.</p> <p>&quot;4xco2&quot; have 4 x pre-industrial atmospheric carbon dioxide concentration. (Others have 1 x pre-industrial atmospheric carbon dioxide concentration.)</p> <p>&quot;rh0.7&quot; and &quot;rh0.9&quot; have LLCS RHCRIT set to 70% and 90% respectively.</p> <p>All are 30 day simulations either for January &quot;jan&quot; or July &quot;jul&quot;.</p> <p>&nbsp;</p>

opencc-by-4.0May 2020View details →
dryad32/100

Data from: The Cumulative Indel Model: fast and accurate statistical evolutionary alignment

Sequence alignment is essential for phylogenetic and molecular evolution inference, as well as in many other areas of bioinformatics and evolutionary biology. Inaccurate alignments can lead to severe biases in most downstream statistical analyses. Statistical alignment based on probabilistic models of sequence evolution addresses these issues by replacing heuristic score functions with evolutionary model-based probabilities. However, score-based aligners and fixed-alignment phylogenetic approaches are still more prevalent than methods based on evolutionary indel models, mostly due to computational convenience. Here, I present new techniques for improving the accuracy and speed of statistical evolutionary alignment. The "cumulative indel model" approximates realistic evolutionary indel dynamics using differential equations. "Adaptive banding" reduces the computational demand of most alignment algorithms without requiring prior knowledge of divergence levels or pseudo-optimal alignments. Using simulations, I show that these methods lead to fast and accurate pairwise alignment inference. Also, I show that it is possible, with these methods, to align and infer evolutionary parameters from a single long synteny block (approximately 530kbp) between the human and chimp genomes. The cumulative indel model and adaptive banding can therefore improve the performance of alignment and phylogenetic methods.

opencc-zeroAug 2020View details →
zenodo32/100

MOSTWAS models, TWAS summary statistics, and simulation results for Bhattacharya and Love, 2020

<p>MOSTWAS models, TWAS results, simulation results, and comparison to BGW-TWAS results</p>

opencc-by-4.0Apr 2020View details →
dryad32/100

Calibration of probability predictions from machine-learning and statistical models

<p>This data set describes the occurrence (yes/no) of a bird, the Southern Whiteface (<i>Aphelocephala leucopsis)</i> in Australia. A suite of environmental variables is provided, which are used in the paper to illustrate a statistical problem. The data are meant to allow reproduction of the analysis in this paper. They are not intended for actual ecological analysis. The data come as .Rdata-file, i.e. as an R-dataset (described technically here: https://www.loc.gov/preservation/digital/formats/fdd/fdd000470.shtml).</p> <p>Here is the paper's abstract:</p> <p><span>Aim: Predictions from statistical models may be uncalibrated, meaning that the predicted values do not have the nominal coverage probability. This is easiest seen with probability predictions in machine-learning classification, including the common species occurrence probabilities. Here, a predicted probability of, say, 0.7 should indicate that out of 100 cases with these environmental conditions, and hence the same predicted probability, the species should be present in 70 and absent in 30.</span><br> <span>Innovation: A simple calibration plot shows that this is not necessarily the case, particularly not for over-fitted models or algorithms that use non-likelihood target functions. As a consequence, "raw" predictions from such model could easily be off by 0.2, are unsuitable for averaging across model types, and resulting maps hence be substantially distorted. The solution, a flexible calibration regression, is simple and can be applied whenever deviations are observed.</span><br> <span>Conclusion: "Raw", uncalibrated probability predictions should be calibrated before interpreting or averaging them in a probabilistic way.</span></p>

opencc-zeroJan 2021View details →
dryad32/100

Data from: Use of simulation-based statistical models to complement bioclimatic models in predicting continental scale invasion risks

Invasive species represent one of the greatest risks to global biodiversity and economic productivity of agroecosystems. The development of certain novel crops—e.g., herbaceous perennial biomass crops—may create a risk of novel invasions by these crops. Therefore, potential benefits and risks need to be weighed in making decisions about their introduction and subsequent management. Ideally, such a weighing will be based on good estimates of invasion risks in realistic scenarios pertaining to actual landscapes of concern regarding invasion. Most previous large-scale analyses of invasion risk have used species distribution models and their established methods. Unfortunately, these approaches are unable to incorporate local scale biotic and spatial factors that influence invasion risk. Here we present a case study for how such factors can be efficiently incorporated in large-scale analyses of invasion risk, by extending simulation models with statistical modeling tools. By these means, we predict invasion risk at the scale of the entire United States for a major biomass crop, Miscanthus × giganteus. We then combine invasion risk predictions for this method with those from bioclimatic methods, producing a map of aggregated invasion risk that can offer more nuanced predictions of invasion risk than either approach alone. Lastly, we evaluate potential risks for invasive crops that differ in invasiveness traits, to examine how geographic patterns of invasion risk vary among invaders as a result of their particular constellation of traits.

opencc-zeroDec 2017View details →
dryad32/100

Data from: A statistical skull geometry model for children 0-3 years old

Head injury is the leading cause of fatality and long-term disability for children. Pediatric heads change rapidly in both size and shape during growth, especially for children under 3 years old (YO). To accurately assess the head injury risks for children, it is necessary to understand the geometry of the pediatric head and how morphologic features influence injury causation within the 0–3 YO population. In this study, head CT scans from fifty-six 0–3 YO children were used to develop a statistical model of pediatric skull geometry. Geometric features important for injury prediction, including skull size and shape, skull thickness and suture width, along with their variations among the sample population, were quantified through a series of image and statistical analyses. The size and shape of the pediatric skull change significantly with age and head circumference. The skull thickness and suture width vary with age, head circumference and location, which will have important effects on skull stiffness and injury prediction. The statistical geometry model developed in this study can provide a geometrical basis for future development of child anthropomorphic test devices and pediatric head finite element models.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Accurate genomic prediction of Coffea canephora in multiple environments using whole-genome statistical models

Genomic selection have been proposed as the standard method to predict breeding values in animal and plant breeding. Although some crops have benefited from this methodology, studies in Coffea are still emerging. To date, there have been no studies of how well genomic prediction models work across populations and environments for different complex traits in coffee. Considering that predictive models are based on biological and statistical assumptions, it is expected that their performance vary depending on how well these assumptions align with the true genetic architecture of the phenotype. To investigate this, we used data from two recurrent selection populations of Coffea canephora, evaluated in two locations, and single nucleotide polymorphisms identified by Genotyping-by-Sequencing. In particular, we evaluated the performance of 13 statistical approaches to predict three important traits in the coffee — production of coffee beans, leaf rust incidence and yield of green beans. Analyses were performed for predictions within-environment, across locations and across populations to assess the reliability of genomic selection. Overall, differences in the prediction accuracy of the competing models were small, although the Bayesian methods showed a modest improvement over other methods, at the cost of more computation time. As expected, predictive accuracy for within-environment analysis, on average, were higher than predictions across locations and across populations. Our results support the potential of genomic selection to reshape traditional plant breeding schemes. In practice, we expect to increase the genetic gain per unit of time by reducing the length cycle of recurrent selection in coffee.

opencc-zeroDec 2017View details →
dryad32/100

Data from: High quality statistical shape modelling of the human nasal cavity and applications

The human nose is a complex organ that shows large morphological variations and has many important functions. However, the relation between shape and function is not yet fully understood. In this work, we present a high quality statistical shape model of the human nose based on clinical CT data of 46 patients. A technique based on cylindrical parametrization was used to create a correspondence between the nasal shapes of the population. Applying principal component analysis on these corresponded nasal cavities resulted in an average nasal geometry and geometrical variations, known as principal components, present in the population with a high precision. The analysis led to 46 principal components, which account for 95 percent of the total geometrical variation captured. These variations are first discussed qualitatively, and the effect on the average nasal shape of the first five principal components is visualized. Hereafter, by using this statistical shape model, two application examples that lead to quantitative data are shown: nasal shape in function of age and gender, and a morphometric analysis of different anatomical regions. Shape models, as the one presented here, can help to get a better understanding of nasal shape and variation, and their relationship with demographic data.

opencc-zeroDec 2017View details →
dryad32/100

Data from: A statistical mechanics framework for constructing non-equilibrium thermodynamic models

<p><span>Far-from-equilibrium phenomena are critical to all natural and engi</span><span>neered systems, and essential to biological processes responsible </span><span>for life. For over a century and a half, since Carnot, Clausius, Maxwell, </span><span>Boltzmann, and Gibbs, among many others, laid the foundation for </span><span>our understanding of equilibrium processes, scientists and engineers </span><span>have dreamed of an analogous treatment of non-equilibrium systems. </span><span>But despite tremendous efforts, a universal theory of non-equilibrium </span><span>behavior akin to equilibrium statistical mechanics and thermodynam</span><span>ics has evaded description. Several methodologies have proved their </span><span>ability to accurately describe complex non-equilibrium systems at </span><span>the macroscopic scale, but their accuracy and predictive capacity is </span><span>predicated on either phenomenological kinetic equations fit to mi</span><span>croscopic data, or on running concurrent simulations at the particle </span><span>level. Instead, we provide a framework for deriving stand-alone macro</span><span>scopic thermodynamics models directly from microscopic physics </span><span>without fitting in overdamped Langevin systems.</span> <span>The only neces</span><span>sary ingredient is a functional form for a parameterized, approximate </span><span>density of states, in analogy to the assumption of a uniform density </span><span>of states in the equilibrium microcanonical ensemble. We highlight </span><span>this framework's effectiveness by deriving analytical approximations </span><span>for evolving mechanical and thermodynamic quantities in a model of </span><span>coiled-coil proteins and double stranded DNA, thus producing, to the </span><span>authors' knowledge, the first derivation of the governing equations for </span><span>a phase propagating system under general loading conditions without </span><span>appeal to phenomenology. The generality of our treatment allows </span><span>for application to any system described by Langevin dynamics with </span><span>arbitrary interaction energies and external driving, including colloidal </span><span>macromolecules, hydrogels, and biopolymers.</span></p>

opencc-zeroNov 2023View details →
zenodo32/100

Resistivity models and EM responses for HBNN and statistical sampling

<p>This dataset is obtained from https://github.com/ersan1234/cageo2021_PhyDLI</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Data for "An Ensemble-Based Statistical Methodology to Detect Differences in Weather and Climate Model Executables" Part 1/2

<p>Ensemble simulations from the weather and climate model COSMO. The data has been used for model verification cases in the corresponding paper (https://doi.org/10.5194/gmd-2021-248).</p> <p>The data is partitioned into the following parts:</p> <ol> <li>gpu_dycore.tar.gz<br> 5-day ensemble (600 members) produced with COSMO 5.09 GPU version in double precision.</li> <li>cpu_nodycore.tar.gz<br> 5-day ensemble (200 members) produced with COSMO 5.09 CPU version in double precision.</li> <li>gpu_dycore_sp.tar.gz<br> 5-day ensemble (200 members) produced with COSMO 5.09 GPU version in single precision.</li> </ol> <p>The second part of the dataset with the diffusion ensembles can be found here: https://doi.org/10.5281/zenodo.6355647</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Dataset for "Modeling Cell Populations Measured By Flow Cytometry With Covariates Using Sparse Mixture of Regressions" in the Annals of Applied Statistics

<p>This is the dataset to be used for the paper in the Annals of Applied Statistics&nbsp;titled:</p> <p><strong>&quot;Modeling Cell Populations Measured By Flow Cytometry With Covariates Using Sparse Mixture of Regressions&quot;</strong></p> <p>Download, unzip and place in the ./<strong>paper-data</strong>&nbsp;directory in the R package&nbsp;repository&nbsp;<a href="https://github.com/sangwon-hyun/flowmix">https://github.com/sangwon-hyun/flowmix</a>. Then, run the code in <strong>./paper-code</strong>&nbsp;to produce the figures and tables.</p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

SEP list and effective dose of paper "The radiation impact of Solar Energetic Particle Events on the Moon: A statistical study using data-based modeling results"

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

Predicting readers' prototypical eye-movement behavior using MASC, a model of Attention in the Superior Colliculus: Stimulus materials, model code, data, and statistical analyses.

<p>The goal of the present research was to determine the role of rudimentary visuo-motor pathways, from the retina and the primary visual cortex to the superior colliculus (SC), in the guidance of human eye movement during reading. To this end, we used MASC, our model of Attention in the Superior Colliculus (Adeli et al., Journal of Neuroscience 2017), a model that relies on well-established saccade-programming principles in the SC. MASC predicts sequences of fixations over an input image by spatially integrating incoming signals in the space of the SC.</p> <p>Here, MASC computed the distribution of luminance contrast over sentences&#39; images (visual-saliency map), after blurring it proportional to retinal eccentricity (retina transformation). It then projected the visual-saliency map into SC space, using a logarithmic afferent-mapping function (magnification factor). Input signals were averaged over retinotopically organized populations of neurons (point images) of constant size, first in the visual map and then in a spatially-registered motor map. The most active population was identified through a winner-take-all process. After jitter applied to the winning population, the next fixation location was determined using inverse efferent mapping. This sequence of events was then repeated to predict following fixation locations, but inserting after each saccade an inhibitory spatial tag (Inhibition of Saccade Return; ISR -referred to as IOR in the uploaded files). All MASC&#39;s parameters, but one, were biologically determined, using electrophysiological data in macaque; the ISR window was the one fit parameter.</p> <p>MASC was tested by comparing its predicted sequences of fixations over sentences from the French-Sentence Corpus (FSC) to the eye-movement behavior of 40 French-native speakers reading the same sentences for comprehension (Albrengues et al., Plos One 2019). Then, MASC was dissected to determine the crucial processing steps enabling prediction of human behavior (10 comparison models -see the general README file). Finally, to address crucial issues in the reading literature, i.e., the role of inter-word spacing and character print size in eye-movement guidance, MASC was additionally tested in four additional display conditions: the same sentences from the FSC, but with blank spaces between words being either filled or removed, or with the screen width angle being multiplied by 2 or 4, such that characters were larger in angular size (0.5&deg; and 1&deg;) than in the original experiment (0.25&deg;). MASC&#39;s predicted effects of inter-word spacing and print size were compared to previously published data.</p> <p>All material relevant to the project is reported here, including the FSC materials (bitmap and information text files), the Matlab code for our MASC model, raw simulation data for MASC and all our comparison models, as well as MASC&#39;s simulations in the different display conditions, the scripts we developed in R to transform raw simulation data into data matrices for statistical analyses of (word-based) eye-movement behavior, the resulting data matrices for all models as well as the data matrix for FSC readers, the R-scripts for statistical comparison of oculomotor behavior between data sets and conditions, literature-review tables of previously published data (for comparison with MASC&#39;s predictions), and the R-scripts generating the figures summarizing our results.</p> <p>Further information can be found in the general README file as well as in the README files attached to each folder. The authors&#39; respective contributions to the project, the licence attached to the included materials and their condition of use are listed in the general README file.</p> <p>A manuscript reporting and discussing these modeling data is in preparation (Vitu, F., Adeli, H. &amp; Zelinsky, G. J.); A reference will be provided here when the manuscript appears in a journal.</p> <p>Other references to be cited:</p> <p>- For the model code: Adeli, H., Vitu, F., &amp; Zelinsky, G. J. (2017). A model of the superior colliculus predicts fixation locations during scene viewing and visual search. Journal of Neuroscience, 37(6), 1453-1467. http://www.jneurosci.org/content/37/6/1453</p> <p>- For FSC materials and data: Albrengues, C., Lavigne, F., Aguilar, C., Castet, E., &amp; Vitu, F. (2019). Linguistic processes do not beat visuo-motor constraints, but they modulate where the eyes move regardless of word boundaries: Evidence against top-down word-based eye-movement control during reading. PLoS ONE 14(7): e0219666. https://doi.org/10.1371/journal.pone.0219666<br> &nbsp;</p>

opencc-by-nc-nd-4.0Aug 2021View details →
zenodo32/100

Forecasting the July Precipitation over the middle-lower reaches of the Yangtze River with a flexible statistical model

<p>The archive.zip contains the MLYR forecast system codes, data, results and all the figures which used in the paper&ldquo;Forecasting the July Precipitation over the middle-lower reaches of the Yangtze River with a flexible statistical model.&rdquo;</p> <ul> <li>The MLYR forecast system codes of each experiment used in this study&nbsp;are under the directory of RegMLYR.</li> <li>The observed data of sea surface temperature and precipitation used in the experiments are under the directory of ObsData.</li> <li>The&nbsp;&nbsp;NCL and Python scripts used for figures in the paper are under the directory of Figuers.</li> <li>The forecasting results of each&nbsp;experiment&nbsp;in this paper are under the directory of Results.</li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Everest Model Output Statistics: improved wind speed forecasts

<p>Data and pre-trained random forest models necessary to correct GFS forecast data and produce improved Everest Forecasts.</p> <p>Please find the associated code at:&nbsp;github.com/MaxVWDV/Everest_wind_forecast</p>

opencc-by-4.0May 2023View details →
zenodo32/100

GAN, PCA, and Statistical Shape Models for the Creation of Synthetic Craniosynostosis Distance Maps

<p>This dataset is part of the publication &quot;Classification of Craniosynostosis Trained Only On Synthetic Data Using GANs, PCA, and Statistical Shape Models&quot;.</p> <p><strong>dataset28.zip</strong> includes 2D distance maps constructed of surface scans of craniosynostosis patients: sagittal suture fusion (scaphocephaly), metopic suture fusion (trigonocephaly), coronal suture fusion (brachycephaly and anterior plagiocephaly), and a control model (normocephaly and positional plagiocephaly).<br> <br> <strong>synthetic_1000.zip</strong> contains are random 1000 samples per class created from each individual synthetic data source (GAN, PCA, statistical shape model).<br> <br> This repository contains only the images. To synthesize your own data, please use the github repository.</p>

opencc-by-nc-4.0Jul 2023View details →
dryad32/100

Data from: Accurate genomic prediction of Coffea canephora in multiple environments using whole-genome statistical models

Open the record for dataset details and reuse information.

publicJun 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record