Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

155

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

155 results for “cluster analysis”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis

The inference of population genetic structures is essential in many research areas in population genetics, conservation biology and evolutionary biology. Recently, unsupervised Bayesian clustering algorithms have been developed to detect a hidden population structure from genotypic data, assuming among others that individuals taken from the population are unrelated. Because of this hypothesis, markers in a sample taken from a subpopulation can be considered to be in Hardy-Weinberg and linkage equilibrium. However, close relatives might be sampled from the same subpopulation, and consequently, might cause Hardy-Weinberg and linkage disequilibrium and thus bias a population genetic structure analysis. In this study, we used simulated and real data to investigate the impact of close relatives in a sample on Bayesian population structure analysis. We also showed that, when close relatives were identified by a pedigree reconstruction approach and removed, the accuracy of a population genetic structure analysis can be greatly improved. The results indicate that unsupervised Bayesian clustering algorithms cannot be used blindly to detect genetic structure in a sample with closely related individuals. Rather, when closely related individuals are suspected to be frequent in a sample, these individuals should be first identified and removed before conducting a population structure analysis.

opencc-zeroDec 2011View details →
zenodo28/100

Data and R code for cluster analysis and machine learning modelling of favourite places for outdoor recreation

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →
zenodo28/100

Spatial Distribution and Cluster Analysis of Road Traffic Accidents in Nepal

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo28/100

Characteristics of individuals with moderate to severe asthma who better respond to aerobic training: a cluster analysis

<p>Database used in the study&#39;s statistical analysis</p>

opencc-by-4.0Mar 2022View details →
zenodo28/100

Dataset for "A Clustering Analysis of Lebanese Adaptive Driving Behaviors in Response to Road Complexity" By Kobeissy et al. Submitted to The Open Transportation Journal

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo28/100

Fig. 2 in Morphometric Analysis And Interrelationship Of Seven Indonesian Hornbill Species (Aves, Bucerotidae) Utilizing Principal Component And Cluster Analysis

Fig. 2. Discriminant function graph of seven hornbill species based on the enter independents together model: A — genus Rhyticeros; B — genus Buceros; C — genus Anthracoceros.

opencc-by-4.0Jun 2024View details →
zenodo28/100

Dataset for Evaluating geopolitical gas supply chain security in the EU: A literature-based index and a clustering analysis

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo28/100

Globular Cluster Fermipy analysis output

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
dryad28/100

The impact of estimator choice: Disagreement in clustering solutions across K estimators for Bayesian analysis of population genetic structure across a wide range of empirical datasets

<p class="CxSpFirst">The software program STRUCTURE is one of the most cited tools for determining population structure. To infer the optimal number of clusters from STRUCTURE output, the Δ<i>K</i> method is often applied. However, a recent study relying on simulated microsatellite data suggested that this method has a downward bias in its estimation of <i>K</i> and is sensitive to uneven sampling. If this finding holds for empirical datasets, conclusions about the scale of gene flow may have to be revised for a large number of studies. To determine the impact of method choice, we applied recently described estimators of <i>K</i> to re-estimate genetic structure in 41 empirical microsatellite datasets; 15 from a broad range of taxa and 26 focused on a diverse phylogenetic group, coral. We compared alternative estimates of <i>K</i> (Puechmaille statistics) with traditional (Δ<i>K</i> and posterior probability) estimates and found widespread disagreement of estimators across datasets. Thus, one estimator alone is insufficient for determining the optimal number of clusters regardless of study organism or evenness of sampling scheme. Subsequent analysis of molecular variance (AMOVA) between clustering solutions did not necessarily clarify which solution was best. To better infer population structure, we suggest a combination of visual inspection of STRUCTURE plots and calculation of the alternative estimators at various thresholds in addition to Δ<i>K</i>. Differences between estimators could reveal patterns with important biological implications, such as the potential for more population structure than previously estimated, as was the case for many studies reanalyzed here.</p>

opencc-zeroOct 2021View details →
zenodo28/100

Supplementary Materials for article on entitled 'Structure and usage do not explain each other: An analysis of German word-initial clusters', published in Linguistics

<p>See the article for a description of the data</p>

opencc-by-4.0Dec 2022View details →
dryad28/100

Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis

Open the record for dataset details and reuse information.

publicJun 2012View details →
dryad28/100

Data from: Graph analysis of cell clusters forming vascular networks

Open the record for dataset details and reuse information.

publicJan 2018View details →
dryad28/100

The impact of estimator choice: Disagreement in clustering solutions across K estimators for Bayesian analysis of population genetic structure across a wide range of empirical datasets

Open the record for dataset details and reuse information.

publicOct 2021View details →
geo24/100

Longitudinal analysis of T-cell receptor repertoires reveals persistence of antigen-driven CD4+ and CD8+ T-cell clusters in Systemic Sclerosis patients

GEO Series GSE156980. Homo sapiens. 24 samples. Type: Other.

openGEO-OpenJan 2021View details →
geo24/100

Genome-scale knockout simulation and clustering analysis of drug-resistant breast cancer cells reveal drug sensitization targets

GEO Series GSE288840. Homo sapiens. 9 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenApr 2025View details →
geo24/100

Deep learning enables accurate clustering with batch effect removal in single-cell RNA-seq analysis

GEO Series GSE146974. Homo sapiens. 3 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMar 2020View details →
geo24/100

Cluster analysis reveals differential transcript profiles associated with resistance training-induced human skeletal muscle hypertrophy

GEO Series GSE42507. Homo sapiens. 44 samples. Type: Expression profiling by array.

openGEO-OpenMay 2013View details →
geo24/100

Bubble-chip analysis of human origin distributions demonstrates on a genomic scale significant clustering into zones and significant association with transcription

GEO Series GSE21110. Homo sapiens. 7 samples. Type: Other.

openGEO-OpenDec 2010View details →
geo24/100

Co-expression analysis reveals gene cluster associated with methylation of enhancers and chromosomal instability under TP63 and TRIM29 regulation [RNA-seq]

GEO Series GSE204811. Homo sapiens. 13 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMar 2024View details →
geo24/100

Transcriptome-Wide Analysis of Human Liver Reveals Age-Related Differences in the Expression of Select Functional Gene Clusters and Evidence for a PPP1R10-Governed ‘Aging Cascade’.

GEO Series GSE183915. Homo sapiens. 17 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record