Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
236
datasets available to search
ShareScore release 0.7.1
Dataset results
236 results for “Modeling Methods”
Induced pluripotent stem cell-derived cardiomyocyte in vitro models: tissue fabrication protocols, assessment methods, and quantitative maturation metrics for benchmarking progress
Open the record for dataset details and reuse information.
Data from: How far can I extrapolate my species distribution model? Exploring Shape, a novel method
Open the record for dataset details and reuse information.
Development of a panel of SNP loci in the emblematic southern damselfly (Coenagrion mercuriale) using a hybrid method: Pitfalls and recommendations for large-scale SNP genotyping in a non-model endangered species
Open the record for dataset details and reuse information.
A fast, precise, in-vivo method for micron-level 3D models of corals using dental scanners
Open the record for dataset details and reuse information.
Data from: A scalable model for simulating multi-round antibody evolution and benchmarking of clonal tree reconstruction methods
Open the record for dataset details and reuse information.
Data for "Function Space Optimization: A symbolic regression method for estimating parameter transfer functions for hydrological models"
<p>This repository contains all geo-physical catchment properties used in the publication "Function Space Optimization: A symbolic regression method for estimating parameter transfer functions for hydrological models".</p>
Statistical model training data for "Continuous Structural Parameterization: A proposed method for representing different model parameterizations within one structure demonstrated for atmospheric convection"
<p>Gzipped CSV files containing convection scheme inputs and outputs used for training.</p> <p>Column format of each file:</p> <p>THETA_IN_1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,Q_IN_1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,DTHETA_1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,DQ_1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28</p> <p>where THETA_IN are input values of potential temperature [K], Q_IN are input values of specific humidity [kg/kg], DTHETA are changes in potential temperature due to convection [K], DQ are changes in specific humidity due to convection [kg/kg].</p> <p>Key:</p> <p>"llcs" are simulations with Lambert-Lewis.</p> <p>"gr" are simulations with Gregory-Rowntree.</p> <p>"4xco2" have 4 x pre-industrial atmospheric carbon dioxide concentration. (Others have 1 x pre-industrial atmospheric carbon dioxide concentration.)</p> <p>"rh0.7" and "rh0.9" have LLCS RHCRIT set to 70% and 90% respectively.</p> <p>All are 30 day simulations either for January "jan" or July "jul".</p> <p> </p>
Data release for 'Ensemble Forecasting of Major Solar Flares: Methods for Combining Models'
<p>This is a release of the data that were used for validation in the paper 'Ensemble Forecasting of Major Solar Flares: Methods for Combining Models' by J. A. Guerra, S. A. Murray, D. S. Bloomfield, and P. T. Gallagher, that has been submitted to the Journal of Space Weather and Space Climate.</p> <p> </p> <ul> </ul> <p>The naming scheme for the files is in the format:</p> <pre><code>class_type_metric.dat</code></pre> <ul> <li>'class' denotes whether the forecast is for M- or X- class flares.</li> <li>'type' is what kind of forecast, i.e., the original ensemble members, an ensemble created from probabilistic validation metrics, or an ensemble created from categorical validation metrics.</li> <li>'metric' specifies the metric used to create the ensemble in the case of 'probabilistic' or 'categorical' types as above (see paper for further details), or in the case of the original ensemble members the name of the operational forecasting method.</li> </ul> <p> </p> <p>The data files are in the format:</p> <pre><code>obs,prob</code></pre> <ul> <li>'obs' denotes whether or not a flare was observed within 24 hours of the forecast issue time (1 for yes and 0 for no).</li> <li>'prob' gives the probabilistic forecast value (between 0.0 and 1.0).</li> </ul> <p> </p> <p>These data files can easily be read into the <a href="https://cran.r-project.org/web/packages/verification/verification.pdf">R verification package</a> to replicate the results presented in the paper.</p>
Data from: A method for mapping morphological convergence on three-dimensional digital models: the case of the mammalian saber-tooth
<p>Morphological convergence can be assessed through a variety of statistical methods. None of the methods proposed to date enable the visualization of convergence. All are based on the assumption that the phenotypes either converge, or do not. However, between species, morphologically similar regions of a larger structure may behave differently. Previous approaches do not identify these regions within the larger structures or quantify the degree to which they may contribute to overall convergence. Here we introduce a new method to chart patterns of convergence on three-dimensional models, deployed with the R function conv.map. The convergence between pairs of models is mapped onto them to visualize and quantify the morphological convergence. We applied conv.map to a well-known case study, the saber-tooth morphotype which has evolved independently among distinct mammalian clades, from placentals to metatherians. Although previous authors have concluded that saber-tooths kill using a stabbing 'bite' to the neck, others have presented different interpretations for specific taxa, including the iconic Smilodon and Thylacosmilus. Our objective was to identify any shared morphological features among the saber-tooths that may underpin similar killing behaviours. From a sample of 49 placental and metatherian carnivores, we found stronger convergence among saber-tooths than for any other taxa. We found that the morphological convergence is most apparent in the rostral and posterior parts of the cranium. The extent of this convergence implies similarity in function among these phylogenetically distant species. In our view this function is most likely the killing of relatively large prey by a stabbing bite.</p>
Data from: Are cranial biomechanical simulation data linked to known diets in extant taxa? A method for applying diet-biomechanics linkage models to infer feeding capability of extinct species
Performance of the masticatory system directly influences feeding and survival, so adaptive hypotheses often are proposed to explain craniodental evolution via functional morphology changes. However, the prevalence of "many-to-one" association of cranial forms and functions in vertebrates suggests a complex interplay of ecological and evolutionary histories, resulting in redundant morphology-diet linkages. Here we examine the link between cranial biomechanical properties for taxa with different dietary preferences in crown clade Carnivora, the most diverse clade of carnivorous mammals. We test whether hypercarnivores and generalists can be distinguished based on cranial mechanical simulation models, and how such diet-biomechanics linkages relate to morphology. Comparative finite element and geometric morphometrics analyses document that predicted bite force is positively allometric relative to skull strain energy; this is achieved in part by increased stiffness in larger skull models and shape changes that resist deformation and displacement. Size-standardized strain energy levels do not reflect feeding preferences; instead, caniform models have higher strain energy than feliform models. This caniform-feliform split is reinforced by a sensitivity analysis using published models for six additional taxa. Nevertheless, combined bite force-strain energy curves distinguish hypercarnivorous versus generalist feeders. These findings indicate that the link between cranial biomechanical properties and carnivoran feeding preference can be clearly defined and characterized, despite phylogenetic and allometric effects. Application of this diet-biomechanics linkage model to an analysis of an extinct stem carnivoramorphan and an outgroup creodont species provides biomechanical evidence for the evolution of taxa into distinct hypercarnivorous and generalist feeding styles prior to the appearance of crown carnivoran clades with similar feeding preferences.
Data from: A model-derived short-term estimation method of effective size for small populations with overlapping generations
If not actively managed, small and isolated populations lose their genetic variability and the inbreeding rate increases. Combined, these factors limit the ability of populations to adapt to environmental changes, increasing their risk of extinction. The effective population size (Ne) is proportional to the loss of genetic diversity and therefore of considerable conservation relevance. However, estimators of Ne that account for demographic parameters in species with overlapping generations require sampling of populations across generations, which is often not feasible in long-lived species. We created an individual-based model that allows calculation of Ne based on demographic parameters that can be obtained in a time period much shorter than a generation. It can be adapted to every life-history parameter combination. The model is freely available as an r-package NEff. The model was first used in a simulation experiment observing changes in Ne in response to different degrees of generational overlap. Results showed that increased generational overlap slowed annual rates of heterozygosity loss, resulting in higher annual effective sizes (Ny) but decreased Ne per generation. Adding the effect of different recruitment rates only affected Ne for populations with low generational overlap. The model was further tested using real population data of the Australian arboreal gecko Gehyra variegata. Simulation results were compared to genetic analyses and matched estimates of the real population very well. Unlike other estimation methods of Ne, NEff neither requires long time series of population monitoring nor genetic analyses of changes in gene frequencies. Thus, it seems to be the first method for calculating Ne within short time periods and comparably low costs facilitating the use of Ne in applied conservation and management.
Data from: Quantifying demographic uncertainty: Bayesian methods for integral projection models (IPMs)
Integral projection models (IPMs) are a powerful and popular approach to modeling population dynamics. Generalized linear models form the statistical backbone of an IPM. These models are typically fit using a frequentist approach. We suggest that hierarchical Bayesian statistical approaches offer important advantages over frequentist methods for building and interpreting IPMs, especially given the hierarchical nature of most demographic studies. Using a stochastic IPM for a desert cactus based on a 10-year study as a worked example, we highlight the application of a Bayesian approach for translating uncertainty in the vital rates (e.g., growth, survival, fertility) to uncertainty in population-level quantities derived from them (e.g., population growth rate). The best-fit demographic model, which would have been difficult to fit under a frequentist framework, allowed for spatial and temporal variation in vital rates and correlated responses to temporal variation across vital rates. The corresponding posterior probability distribution for the stochastic population growth rate (λS) indicated that, if current vital rates continue, the study population will decline with nearly 100% probability. Interestingly, less-supported candidate models that did not include spatial variance and vital rate correlations gave similar estimates of λS. This occurred because the best-fitting model did a much better job of fitting vital rates to which the population growth rate was weakly sensitive. The cactus case study highlights several advantages of Bayesian approaches to IPM modeling, including that they: (1) provide a natural fit to demographic data, which are often collected in a hierarchical fashion (e.g., with random variance corresponding to temporal and spatial heterogeneity); (2) seamlessly combine multiple data sets or experiments; (3) readily incorporate covariance between vital rates; and, (4) easily integrate prior information, which may be particularly important for species of conservation concern where data availability may be limited. However, constructing a Bayesian IPM will often require the custom development of a statistical model tailored to the peculiarities of the sampling design and species considered; there may be circumstances under which simpler methods are adequate. Overall, Bayesian approaches provide a statistically sound way to get more information out of hard-won data, the goal of most demographic research endeavors.
Data from: Linkage into care among newly diagnosed HIV-positive individuals tested through outreach and facility-based HIV testing models in Mbeya, Tanzania: a prospective mixed-method cohort study
Objective: Linkage to care is the bridge between HIV testing and HIV treatment, care and support. In Tanzania, mobile testing aims to address historically low testing rates. Linkage to care was reported at 14% in 2009 and 28% in 2014. The study compares linkage to care of HIV-positive individuals tested at mobile/outreach versus public health facility-based services within the first 6 months of HIV diagnosis. Setting: Rural communities in four districts of Mbeya Region, Tanzania. Participants: A total of 1012 newly diagnosed HIV-positive adults from 16 testing facilities were enrolled into a two-armed cohort and followed for 6 months between August 2014 and July 2015. 840 (83%) participants completed the study. Main outcome: measures We compared the ratios and time variance in linkage to care using the Kaplan-Meier estimator and Log rank tests. Cox proportional hazards regression models to evaluate factors associated with time variance in linkage. Results: At the end of 6 months, 78% of all respondents had linked into care, with differences across testing models. 84% (CI 81% to 87%, n=512) of individuals tested at facility-based site were linked to care compared to 69% (CI 65% to 74%, n=281) of individuals tested at mobile/outreach. The median time to linkage was 1 day (IQR: 1–7.5) for facility-based site and 6 days (IQR: 3–11) for mobile/outreach sites. Participants tested at facility-based site were 78% more likely to link than those tested at mobile/outreach when other variables were controlled (AHR=1.78; 95% CI 1.52 to 2.07). HIV status disclosure to family/relatives was significantly associated with linkage to care (AHR=2.64; 95% CI 2.05 to 3.39). Conclusions: Linkage to care after testing HIV positive in rural Tanzania has increased markedly since 2014, across testing models. Individuals tested at facility-based sites linked in significantly higher proportion and modestly sooner than mobile/outreach tested individuals. Mobile/outreach testing models bring HIV testing services closer to people. Strategies to improve linkage from mobile/outreach models are needed.
Data from: A rapid and scalable method for multilocus species delimitation using Bayesian model comparison and rooted triplets
Multilocus sequence data provide far greater power to resolve species limits than the single locus data typically used for broad surveys of clades. However, current statistical methods based on a multispecies coalescent framework are computationally demanding, because of the number of possible delimitations that must be compared and time-consuming likelihood calculations. New methods are therefore needed to open up the power of multilocus approaches to larger systematic surveys. Here, we present a rapid and scalable method that introduces 2 new innovations. First, the method reduces the complexity of likelihood calculations by decomposing the tree into rooted triplets. The distribution of topologies for a triplet across multiple loci has a uniform trinomial distribution when the 3 individuals belong to the same species, but a skewed distribution if they belong to separate species with a form that is specified by the multispecies coalescent. A Bayesian model comparison framework was developed and the best delimitation found by comparing the product of posterior probabilities of all triplets. The second innovation is a new dynamic programming algorithm for finding the optimum delimitation from all those compatible with a guide tree by successively analyzing subtrees defined by each node. This algorithm removes the need for heuristic searches used by current methods, and guarantees that the best solution is found and potentially could be used in other systematic applications. We assessed the performance of the method with simulated, published, and newly generated data. Analyses of simulated data demonstrate that the combined method has favorable statistical properties and scalability with increasing sample sizes. Analyses of empirical data from both eukaryotes and prokaryotes demonstrate its potential for delimiting species in real cases.
Data from: A comparison of genomic selection models across time in interior spruce (Picea engelmannii × glauca) using unordered SNP imputation methods
Genomic selection (GS) potentially offers an unparalleled advantage over traditional pedigree-based selection (TS) methods by reducing the time commitment required to carry out a single cycle of tree improvement. This quality is particularly appealing to tree breeders, where lengthy improvement cycles are the norm. We explored the prospect of implementing GS for interior spruce (Picea engelmannii × glauca) utilizing a genotyped population of 769 trees belonging to 25 open-pollinated families. A series of repeated tree height measurements through ages 3–40 years permitted the testing of GS methods temporally. The genotyping-by-sequencing (GBS) platform was used for single nucleotide polymorphism (SNP) discovery in conjunction with three unordered imputation methods applied to a data set with 60% missing information. Further, three diverse GS models were evaluated based on predictive accuracy (PA), and their marker effects. Moderate levels of PA (0.31–0.55) were observed and were of sufficient capacity to deliver improved selection response over TS. Additionally, PA varied substantially through time accordingly with spatial competition among trees. As expected, temporal PA was well correlated with age-age genetic correlation (r=0.99), and decreased substantially with increasing difference in age between the training and validation populations (0.04–0.47). Moreover, our imputation comparisons indicate that k-nearest neighbor and singular value decomposition yielded a greater number of SNPs and gave higher predictive accuracies than imputing with the mean. Furthermore, the ridge regression (rrBLUP) and BayesCπ (BCπ) models both yielded equal, and better PA than the generalized ridge regression heteroscedastic effect model for the traits evaluated.
Deep Learning Methods for Unsupervised Acoustic Modeling using HMM posteriograms (system #2)
<p>System combination of HMM-DNN with auto encoder features</p>
Deep Learning Methods for Unsupervised Acoustic Modeling using GMM posteriograms (system #1)
<p>System combination of autoencoder and GMM-DNN features. </p>
Deep Learning Methods for Unsupervised Acoustic Modeling using HMM posteriograms
<p>DNN trained using HMM posteriograms</p>
Dataset for a physics informed deep learning method with adaptively weighted loss for modeling soil water flows
<p>The data for the 11 scenarios generated by Hydrus-1D is located in data.zip</p> <p>The code for the physics-informed neural networks with adaptively weighted loss used to simulate water flow in loam soils is located at PINN_adaptively_weighted_loss_loam.zip</p>
Dataset associated with "Emulating subglacial hydrology in ice sheet models with deep learning methods" by Verjans and Robel.
<p>See Readme file for descriptions.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.