Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
94
datasets available to search
ShareScore release 0.9.0
Dataset results
94 results for “methods: data analysis”
Data from: A method for analysis of phenotypic change for phenotypes described by high-dimensional data
Open the record for dataset details and reuse information.
Data from: Is computer-assisted instruction more effective than other educational methods in achieving ECG competence amongst medical students and residents? A systematic review and meta-analysis.
Open the record for dataset details and reuse information.
Data from: Free vibration analysis of transmission lines based on the dynamic stiffness method
Open the record for dataset details and reuse information.
Data from: Integrating complementary methods to improve diet analysis in fishery-targeted species
Open the record for dataset details and reuse information.
Data from: Comparative analysis of DNA extraction methods to study the body surface microbiota of insects: a case study with ant cuticular bacteria
Open the record for dataset details and reuse information.
Data from: Identifying conservation priorities for gorgonian forests in Italian coastal waters with multiple methods including citizen science and social media content analysis
Open the record for dataset details and reuse information.
Data from: A new method for studying population genetics of cyst nematodes based on Pool-Seq and genome-wide allele frequency analysis
Open the record for dataset details and reuse information.
Data from: Survival analysis and classification methods for forest fire size
Open the record for dataset details and reuse information.
Data from: Extracting spatio-temporal patterns in animal trajectories: an ecological application of sequence analysis methods
Open the record for dataset details and reuse information.
Data for: Theropod dinosaur diversity of the lower English Wealden: analysis of a tooth-based fauna from the Wadhurst Clay Formation (Lower Cretaceous: Valanginian) via phylogenetic, discriminant and machine learning methods
Open the record for dataset details and reuse information.
Data from: Making soil particle size analysis by laser diffraction compatible with standard soil texture determination methods
Open the record for dataset details and reuse information.
Data from: A comparison of regression methods for model selection in individual-based landscape genetic analysis
Open the record for dataset details and reuse information.
Data from: Convergence of multiple markers and analysis methods defines the genetic distinctiveness of cryptic pitvipers
Open the record for dataset details and reuse information.
Data from: A comparative analysis of common methods to identify waterbird hotspots
Open the record for dataset details and reuse information.
Data associated with "Adjoint methods for stellarator shape optimization and sensitivity analysis"
<p>All data produced for this dissertation and the associated post-processing scripts have been archived. </p>
Data from: CLIP test: a new fast, simple and powerful method to distinguish between linked or pleiotropic quantitative trait loci in linkage disequilibria analysis
An important question arises when mapping quantitative trait loci (QTLs) for genetically correlated traits: is the correlation due to pleiotropy (a single QTL affecting more than one trait) and/or close linkage (different QTLs that are physically close to each other and influence the traits)? In this article, we propose the Close Linkage versus Pleiotropism (CLIP) test, a fast, simple and powerful method to distinguish between these two situations. The CLIP test is based on the comparison of the square of the observed correlation between a combination of apparent effects at the marker level to the minimal value it can take under the pleiotropic assumption. A simulation study was performed to estimate the power and alpha risk of the CLIP test and compare it to a test that evaluated whether the confidence intervals of the two QTLs overlapped or not (CI test). On average, the CLIP test showed a higher power (68%) to detect close-linked QTLs than the CI test (43%) and a same alpha risk (4%).
Data from: High throughput method for analysis of repeat number for 28 phase variable loci of Campylobacter jejuni strain NCTC11168
Mutations in simple sequence repeat tracts are a major mechanism of phase variation in several bacterial species including Campylobacter jejuni. Changes in repeat number of tracts located within the reading frame can produce a high frequency of reversible switches in gene expression between ON and OFF states. The genome of C. jejuni strain NCTC11168 contains 29 loci with polyG/polyC tracts of seven or more repeats. This protocol outlines a method—the 28-locus-CJ11168 PV-analysis assay—for rapidly determining ON/OFF states of 28 of these phase-variable loci in a large number of individual colonies from C. jejuni strain NCTC11168. The method combines a series of multiplex PCR assays with a fragment analysis assay and automated extraction of fragment length, repeat number and expression state. This high throughput, multiplex assay has utility for detecting shifts in phase variation states within and between populations over time and for exploring the effects of phase variation on adaptation to differing selective pressures. Application of this method to analysis of the 28 polyG/polyC tracts in 90 C. jejuni colonies detected a 2.5-fold increase in slippage products as tracts lengthened from G8 to G11 but no difference between tracts of similar length indicating that flanking sequence does not influence slippage rates. Comparison of this observed slippage to previously measured mutation rates for G8 and G11 tracts in C. jejuni indicates that PCR amplification of a DNA sample will over-estimate phase variation frequencies by 20-35-fold. An important output of the 28-locus-CJ11168 PV-analysis assay is combinatorial expression states that cannot be determined by other methods. This method can be adapted to analysis of phase variation in other C. jejuni strains and in a diverse range of bacterial species.
Data from: Relative accuracy of three common methods of parentage analysis in natural populations
Parentage studies and family reconstructions have become increasingly popular for investigating a range of evolutionary, ecological and behavioral processes in natural populations. However, a number of different assignment methods have emerged in common use, and the accuracy of each may differ in relation to the number of loci examined, allelic diversity, incomplete sampling of all candidate parents, and the presence of genotyping errors. Here we examine how these factors affect the accuracy of three popular parentage inference methods (COLONY, FaMoz and an exclusion-Bayes' theorem approach by Christie et al. (2010a)) to resolve true parent-offspring pairs using simulated data. Our findings demonstrate that accuracy increases with the number and diversity of loci. These were clearly the most important factors in obtaining accurate assignments explaining 75-90% of variance in overall accuracy across 60 simulated scenarios. Furthermore, the proportion of candidate parents sampled had a small but significant impact on the susceptibility of each method to either false positive or false negative assignments. Within the range of values simulated, COLONY outperformed FaMoz, which outperformed the exclusion-Bayes' theorem method. However, with 20 or more highly polymorphic loci, all methods could be applied with confidence. Our results show that for parentage inference in natural populations, careful consideration of the number and quality of markers will increase the accuracy of assignments and mitigate the effects of incomplete sampling of parental populations.
Data from: Analysis of a rapid evolutionary radiation using ultraconserved elements (UCEs): Evidence for a bias in some multi-species coalescent methods
Rapid evolutionary radiations are expected to require large amounts of sequence data to resolve. To resolve these types of relationships many systematists believe that it will be necessary to collect data by next-generation sequencing (NGS) and use multispecies coalescent ("species tree") methods. Ultraconserved element (UCE) sequence capture is becoming a popular method to leverage the high throughput of NGS to address problems in vertebrate phylogenetics. Here we examine the performance of UCE data for gallopheasants (true pheasants and allies), a clade that underwent a rapid radiation 10–15 Ma. Relationships among gallopheasant genera have been difficult to establish. We used this rapid radiation to assess the performance of species tree methods, using ∼600 kilobases of DNA sequence data from ∼1500 UCEs. We also integrated information from traditional markers (nuclear intron data from 15 loci and three mitochondrial gene regions). Species tree methods exhibited troubling behavior. Two methods [Maximum Pseudolikelihood for Estimating Species Trees (MP-EST) and Accurate Species TRee ALgorithm (ASTRAL)] appeared to perform optimally when the set of input gene trees was limited to the most variable UCEs, though ASTRAL appeared to be more robust than MP-EST to input trees generated using less variable UCEs. In contrast, the rooted triplet consensus method implemented in Triplec performed better when the largest set of input gene trees was used. We also found that all three species tree methods exhibited a surprising degree of dependence on the program used to estimate input gene trees, suggesting that the details of likelihood calculations (e.g., numerical optimization) are important for loci with limited phylogenetic information. As an alternative to summary species tree methods we explored the performance of SuperMatrix Rooted Triple - Maximum Likelihood (SMRT-ML), a concatenation method that is consistent even when gene trees exhibit topological differences due to the multispecies coalescent. We found that SMRT-ML performed well for UCE data. Our results suggest that UCE data have excellent prospects for the resolution of difficult evolutionary radiations, though specific attention may need to be given to the details of the methods used to estimate species trees.
Data from: Testing Phylogenetic Methods with Tree Congruence: Phylogenetic Analysis of Polymorphic Morphological Characters in Phrynosomatid Lizards
Congruence between trees from separately analyzed data sets is a powerful approach for assessing the performance of phylogenetic methods but has been applied primarily to the analysis of molecular data. In this study, different methods for treating polymorphic characters were compared using morphological data from phrynosomatid lizards. Clades were identified that are both traditionally recognized and supported by recent molecular analyses, and species were sampled from these clades to make three RknownS phylogenies of eight species each. The ability of different methods to estimate these "known" phylogenies with a finite sample of characters was tested. The phylogenetic methods included eight parsimony methods for coding polymorphism, three distance approaches (UPGMA, neighbor joining, and Fitch-Margoliash) applied to two genetic distance measures (Nei's and the modified Cavalli-Sforza and Edwards chord distance), and continuous maximum likelihood. The effects of excluding polymorphic characters and character weighting (a priori and successive) were also tested. Among the different parsimony approaches, the fixed-only method (excluding all polymorphic characters) performed relatively poorly, whereas the frequency method (including all polymorphic characters) performed relatively well. However, frequency-based distance methods consistently outperformed parsimony, especially with a small sample size (n= 1 individual per species). These results agree closely with those from recent simulation studies of polymorphic data and argue against the common practices of excluding polymorphic morphological characters, ignoring the frequencies of traits within species, and the exclusive use of parsimony to analyze morphological data.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.