Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
46
datasets available to search
ShareScore release 0.9.0
Dataset results
46 results for “statistical methods”
A statistical method to optimize the chemical etching process of Zinc Oxide thin films
<p>Zinc Oxide (ZnO) is an attractive material for micro and nanoscale devices. Its desirable semiconductor, piezoelectric, and optical properties <span>make</span> it useful in applications ranging from microphones to missile warning systems to biometric sensors. This work introduces a demonstration of blending statistics and chemical etching of thin films to identify the dominant factors, and interaction between factors, and <span>develop</span> statistically enhanced models on etch rate and selectivity of ZnO thin films. Over other mineral acids, ammonium chloride (NH<sub>4</sub>Cl) solutions <span>have</span> commonly been used to <span>wet</span> etch microscale ZnO devices because of their controllable etch rate and near-linear behavior. <span>Etchant</span> concentration and temperature were found to <span>have</span> a significant effect on etch rate. Moreover, this is the first demonstration that has identified multifactor interactions between temperature and concentration and between temperature and agitation. A linear model <span>was</span> developed relating etch rate and its variance against these significant factors and multifactor interactions. An average selectivity of 73:1 <span>was </span>measured with none of the experimental factors having a significant effect on the <span>selectivity. </span>This statistical study captures the significant variance observed <span>by</span> other researchers. <span>Furthermore,</span> it enables statistically enhanced microfabrication processes for other materials.</p>
Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets
The estimation of multiple sequence alignments of protein sequences is a basic step in many bioinformatics pipelines, including protein structure prediction, protein family identification, and phylogeny estimation. Statistical co-estimation of alignments and trees under stochastic models of sequence evolution has long been considered the most rigorous technique for estimating alignments and trees, but little is known about the accuracy of such methods on biological benchmarks. We report the results of an extensive study evaluating the most popular protein alignment methods as well as the statistical co-estimation method BAli-Phy on 1192 protein data sets from established benchmarks as well as on 120 simulated data sets. Our study (which used more than 230 CPU years for the BAli-Phy analyses alone) shows that BAli-Phy has better precision and recall (with respect to the true alignments) than the other alignment methods on the simulated data sets, but has consistently lower recall on the biological benchmarks (with respect to the reference alignments) than many of the other methods. In other words, we find that BAli-Phy systematically under-aligns when operating on biological sequence data, but shows no sign of this on simulated data. There are several potential causes for this change in performance, including model misspecification, errors in the reference alignments, and conflicts between structural alignment and evolutionary alignments, and future research is needed to determine the most likely explanation. We conclude with a discussion of the potential ramifications for each of these possibilities.
Data and code: High-resolution CMIP6 climate projections for Ethiopia using the gridded statistical downscaling method
<p>Data and code supporting the research article:High-resolution CMIP6 climate projections for Ethiopia using the gridded statistical downscaling method - <br> Fasil M. Rettie, Sebastian Gayler, Tobias KD Weber, Kindie Tesfaye, Thilo Streck. Please, find detail description of the codes and datasets in readme file.</p>
A statistical method to optimize the chemical etching process of Zinc Oxide thin films
Open the record for dataset details and reuse information.
Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets
Open the record for dataset details and reuse information.
Data from: A new method for statistical detection of directional and stabilizing mating preference
Open the record for dataset details and reuse information.
Data from: The relative power of genome scans to detect local adaptation depends on sampling design and statistical method
Although genome scans have become a popular approach towards understanding the genetic basis of local adaptation, the field still does not have a firm grasp on how sampling design and demographic history affect the performance of genome scans on complex landscapes. To explore these issues, we compared 20 different sampling designs in equilibrium (i.e. island model and isolation by distance) and nonequilibrium (i.e. range expansion from one or two refugia) demographic histories in spatially heterogeneous environments. We simulated spatially complex landscapes, which allowed us to exploit local maxima and minima in the environment in 'pair' and 'transect' sampling strategies. We compared FST outlier and genetic–environment association (GEA) methods for each of two approaches that control for population structure: with a covariance matrix or with latent factors. We show that while the relative power of two methods in the same category (FST or GEA) depended largely on the number of individuals sampled, overall GEA tests had higher power in the island model and FST had higher power under isolation by distance. In the refugia models, however, these methods varied in their power to detect local adaptation at weakly selected loci. At weakly selected loci, paired sampling designs had equal or higher power than transect or random designs to detect local adaptation. Our results can inform sampling designs for studies of local adaptation and have important implications for the interpretation of genome scans based on landscape data.
Data from: Quantification and statistical analysis methods for vessel wall components from stained images with Masson's trichrome
Purpose: To develop a digital image processing method to quantify structural components (smooth muscle fibers and extracellular matrix) in the vessel wall stained with Masson's trichrome, and a statistical method suitable for small sample sizes to analyze the results previously obtained. Methods: The quantification method comprises two stages. The pre-processing stage improves tissue image appearance and the vessel wall area is delimited. In the feature extraction stage, the vessel wall components are segmented by grouping pixels with a similar color. The area of each component is calculated by normalizing the number of pixels of each group by the vessel wall area. Statistical analyses are implemented by permutation tests, based on resampling without replacement from the set of the observed data to obtain a sampling distribution of an estimator. The implementation can be parallelized on a multicore machine to reduce execution time. Results: The methods have been tested on 48 vessel wall samples of the internal saphenous vein stained with Masson's trichrome. The results show that the segmented areas are consistent with the perception of a team of doctors and demonstrate good correlation between the expert judgments and the measured parameters for evaluating vessel wall changes. Conclusion: The proposed methodology offers a powerful tool to quantify some components of the vessel wall. It is more objective, sensitive and accurate than the biochemical and qualitative methods traditionally used. The permutation tests are suitable statistical techniques to analyze the numerical measurements obtained when the underlying assumptions of the other statistical techniques are not met.
Data from: Computational performance and statistical accuracy of *BEAST and comparisons with other methods
Under the multispecies coalescent model of molecular evolution, gene trees have independent evolutionary histories within a shared species tree. In comparison, supermatrix concatenation methods assume that gene trees share a single common genealogical history, thereby equating gene coalescence with species divergence. The multispecies coalescent is supported by previous studies which found that its predicted distributions fit empirical data, and that concatenation is not a consistent estimator of the species tree. *BEAST, a fully Bayesian implementation of the multispecies coalescent, is popular but computationally intensive, so the increasing size of phylogenetic data sets is both a computational challenge and an opportunity for better systematics. Using simulation studies, we characterize the scaling behavior of *BEAST, and enable quantitative prediction of the impact increasing the number of loci has on both computational performance and statistical accuracy. Follow-up simulations over a wide range of parameters show that the statistical performance of *BEAST relative to concatenation improves both as branch length is reduced and as the number of loci is increased. Finally, using simulations based on estimated parameters from two phylogenomic data sets, we compare the performance of a range of species tree and concatenation methods to show that using *BEAST with tens of loci can be preferable to using concatenation with thousands of loci. Our results provide insight into the practicalities of Bayesian species tree estimation, the number of loci required to obtain a given level of accuracy and the situations in which supermatrix or summary methods will be outperformed by the fully Bayesian multispecies coalescent.
Statistical analysis and dataset for: Linepithema humile shows a lower drinking acceptance for two psychoactive chemicals when using a novel dual-feeder method
Open the record for dataset details and reuse information.
The effect of different statistical methods on the accuracy of predicting genomic selection in beef cattle
Open the record for dataset details and reuse information.
WAYS TO USE THE SAMPLING OBSERVATION METHOD IN THE STATISTICAL ASSESSMENT OF UTILITY ACTIVITIES
Open the record for dataset details and reuse information.
TPP statistics data for: Plurifaceted proteomics method identifies a key regulator of translation during stem cells maintenance and differentiation
<p>Statistics data from TPP package.</p>
Figure 5 in Taxonomic notes on the Asian frogs of the tribe Paini (Ranidae, Dicroglossinae): 1. Morphology and synonymy of Chaparana aenea (Smith, 1922), with proposal of a new statistical method for testing homogeneity of small samples
Figure 5. Plots of factors 1 and 2 of principal component analysis based on varimax rotated coefficients for logtransposed characters (32 measurements) for 13 specimens of Chaparana aenea (Smith, 1922) from China, Thailand and Vietnam (see Table II).
Figure 2 in Taxonomic notes on the Asian frogs of the tribe Paini (Ranidae, Dicroglossinae): 1. Morphology and synonymy of Chaparana aenea (Smith, 1922), with proposal of a new statistical method for testing homogeneity of small samples
Figure 2. Map showing the localities of specimens studied of the species Chaparana aenea (Smith, 1922) (circles) and Chaparana unculuanus (Liu, Hu and Yang, 1960) (square). 1, Doi Chang, Thailand; 2, Doi Inthanon, Thailand; 3, Fan Si Pan, Vietnam; 4, Pu Hoat, Vietnam; 5, Luchun, China; 6, Jingdong, China.
Forecasting ED Overcrowding With Statistical Methods: A Prospective Validation Study
ClinicalTrials.gov study NCT05174481. IPD Sharing: NO. Countries: 0. Publications: 2.
Association Between Mother and Child Weight Gain: Statistical Methods With Validation
ClinicalTrials.gov study NCT04174521. IPD Sharing: NO. Countries: 1. Publications: 0.
Data from: Computational performance and statistical accuracy of *BEAST and comparisons with other methods
Open the record for dataset details and reuse information.
Data from: Mie scattering and microparticle based characterization of heavy metal ions and classification by statistical inference methods
Open the record for dataset details and reuse information.
Data from: Quantification and statistical analysis methods for vessel wall components from stained images with Masson's trichrome
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.