Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
155
datasets available to search
ShareScore release 0.9.0
Dataset results
155 results for “cluster analysis”
Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis
The inference of population genetic structures is essential in many research areas in population genetics, conservation biology and evolutionary biology. Recently, unsupervised Bayesian clustering algorithms have been developed to detect a hidden population structure from genotypic data, assuming among others that individuals taken from the population are unrelated. Because of this hypothesis, markers in a sample taken from a subpopulation can be considered to be in Hardy-Weinberg and linkage equilibrium. However, close relatives might be sampled from the same subpopulation, and consequently, might cause Hardy-Weinberg and linkage disequilibrium and thus bias a population genetic structure analysis. In this study, we used simulated and real data to investigate the impact of close relatives in a sample on Bayesian population structure analysis. We also showed that, when close relatives were identified by a pedigree reconstruction approach and removed, the accuracy of a population genetic structure analysis can be greatly improved. The results indicate that unsupervised Bayesian clustering algorithms cannot be used blindly to detect genetic structure in a sample with closely related individuals. Rather, when closely related individuals are suspected to be frequent in a sample, these individuals should be first identified and removed before conducting a population structure analysis.
Data and R code for cluster analysis and machine learning modelling of favourite places for outdoor recreation
Open the record for dataset details and reuse information.
Spatial Distribution and Cluster Analysis of Road Traffic Accidents in Nepal
Open the record for dataset details and reuse information.
Characteristics of individuals with moderate to severe asthma who better respond to aerobic training: a cluster analysis
<p>Database used in the study's statistical analysis</p>
Dataset for "A Clustering Analysis of Lebanese Adaptive Driving Behaviors in Response to Road Complexity" By Kobeissy et al. Submitted to The Open Transportation Journal
Open the record for dataset details and reuse information.
Fig. 2 in Morphometric Analysis And Interrelationship Of Seven Indonesian Hornbill Species (Aves, Bucerotidae) Utilizing Principal Component And Cluster Analysis
Fig. 2. Discriminant function graph of seven hornbill species based on the enter independents together model: A — genus Rhyticeros; B — genus Buceros; C — genus Anthracoceros.
Dataset for Evaluating geopolitical gas supply chain security in the EU: A literature-based index and a clustering analysis
Open the record for dataset details and reuse information.
Globular Cluster Fermipy analysis output
Open the record for dataset details and reuse information.
The impact of estimator choice: Disagreement in clustering solutions across K estimators for Bayesian analysis of population genetic structure across a wide range of empirical datasets
<p class="CxSpFirst">The software program STRUCTURE is one of the most cited tools for determining population structure. To infer the optimal number of clusters from STRUCTURE output, the Δ<i>K</i> method is often applied. However, a recent study relying on simulated microsatellite data suggested that this method has a downward bias in its estimation of <i>K</i> and is sensitive to uneven sampling. If this finding holds for empirical datasets, conclusions about the scale of gene flow may have to be revised for a large number of studies. To determine the impact of method choice, we applied recently described estimators of <i>K</i> to re-estimate genetic structure in 41 empirical microsatellite datasets; 15 from a broad range of taxa and 26 focused on a diverse phylogenetic group, coral. We compared alternative estimates of <i>K</i> (Puechmaille statistics) with traditional (Δ<i>K</i> and posterior probability) estimates and found widespread disagreement of estimators across datasets. Thus, one estimator alone is insufficient for determining the optimal number of clusters regardless of study organism or evenness of sampling scheme. Subsequent analysis of molecular variance (AMOVA) between clustering solutions did not necessarily clarify which solution was best. To better infer population structure, we suggest a combination of visual inspection of STRUCTURE plots and calculation of the alternative estimators at various thresholds in addition to Δ<i>K</i>. Differences between estimators could reveal patterns with important biological implications, such as the potential for more population structure than previously estimated, as was the case for many studies reanalyzed here.</p>
Supplementary Materials for article on entitled 'Structure and usage do not explain each other: An analysis of German word-initial clusters', published in Linguistics
<p>See the article for a description of the data</p>
Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis
Open the record for dataset details and reuse information.
Data from: Graph analysis of cell clusters forming vascular networks
Open the record for dataset details and reuse information.
The impact of estimator choice: Disagreement in clustering solutions across K estimators for Bayesian analysis of population genetic structure across a wide range of empirical datasets
Open the record for dataset details and reuse information.
Longitudinal analysis of T-cell receptor repertoires reveals persistence of antigen-driven CD4+ and CD8+ T-cell clusters in Systemic Sclerosis patients
GEO Series GSE156980. Homo sapiens. 24 samples. Type: Other.
Genome-scale knockout simulation and clustering analysis of drug-resistant breast cancer cells reveal drug sensitization targets
GEO Series GSE288840. Homo sapiens. 9 samples. Type: Expression profiling by high throughput sequencing.
Deep learning enables accurate clustering with batch effect removal in single-cell RNA-seq analysis
GEO Series GSE146974. Homo sapiens. 3 samples. Type: Expression profiling by high throughput sequencing.
Cluster analysis reveals differential transcript profiles associated with resistance training-induced human skeletal muscle hypertrophy
GEO Series GSE42507. Homo sapiens. 44 samples. Type: Expression profiling by array.
Bubble-chip analysis of human origin distributions demonstrates on a genomic scale significant clustering into zones and significant association with transcription
GEO Series GSE21110. Homo sapiens. 7 samples. Type: Other.
Co-expression analysis reveals gene cluster associated with methylation of enhancers and chromosomal instability under TP63 and TRIM29 regulation [RNA-seq]
GEO Series GSE204811. Homo sapiens. 13 samples. Type: Expression profiling by high throughput sequencing.
Transcriptome-Wide Analysis of Human Liver Reveals Age-Related Differences in the Expression of Select Functional Gene Clusters and Evidence for a PPP1R10-Governed ‘Aging Cascade’.
GEO Series GSE183915. Homo sapiens. 17 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.