Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

23

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

23 results for “Bayesian clustering”

Learn how ShareScore rates datasets ↗
zenodo40/100

Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: Calibration Data

<p>Calibration data accompanying our work, &quot;Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics&quot; by J Bryan IV, I Sgouralis, and S Presse.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 20 Binding Site Data A

<p>This is the original data for the manuscript &quot;Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics&quot; by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 20 binding sites. Because this data set is too large to fit in one single repository we have split it up into parts. This is part A</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 20 Binding Site Data C

<p>This is the original data for the manuscript &quot;Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics&quot; by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 20 binding sites. Because this data set is too large to fit in one single repository we have split it up into parts. This is part C.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 20 Binding Site Data B

<p>This is the original data for the manuscript &quot;Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics&quot; by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 20 binding sites. Because this data set is too large to fit in one single repository we have split it up into parts. This is part B.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Bayesian Samples and Data Behind Figures: Comprehensive Bayesian Modeling of Tidal Circularization in Open Cluster Binaries part I

<p>Auxiliary data associated with the article <a href="https://ui.adsabs.harvard.edu/abs/2022MNRAS.516.6145P/abstract">&quot;Comprehensive Bayesian Modeling of Tidal Circularization in Open Cluster Binaries part I: M 35, NGC 6819, NGC 188&quot; by Penev, K &amp; Schussler, J</a></p> <p>The type of data corresponds to a particular filename format. Bayesian samples are in HDF5 format, directly as saved by the <a href="https://emcee.readthedocs.io/en/stable/index.html">emcee</a> sampler (see <a href="https://emcee.readthedocs.io/en/stable/user/backends/">https://emcee.readthedocs.io/en/stable/user/backends/</a>). All other files are in AAS-journal style machine readable tables format generated by <a href="https://github.com/cds-astro/cds.pyreadme">cdspyreadme</a> python library.</p> <p>Description of contents by filename format:</p> <pre><code>&lt;CLUSTER&gt;_&lt;BINARY ID&gt;_.*.h5</code></pre> <p>Bayesian analysis samples constraining the tidal dissipation efficiency of the given binary. The values of the sampled system and tidal dissipation parameters are stored as blobs (<a href="https://emcee.readthedocs.io/en/stable/user/blobs/">https://emcee.readthedocs.io/en/stable/user/blobs/)</a></p> <pre><code>&lt;CLUSTER&gt;_&lt;BINARY ID&gt;_lgQ_period.mrt</code></pre> <p>The 2.3%, 15.9%, 84.1%, and 97.7% quantiles of <span class="math-tex">\(\log_{10}Q_\star'\)</span> for the given binary as a function of tidal period</p> <pre><code>&lt;CLUSTER&gt;_&lt;BINARY ID&gt;_burnin_period.mrt</code></pre> <p>The MCMC burn-in period before the 2.3%, 15.9%, 84.1%, and 97.7% quantiles of <span class="math-tex">\(\log_{10}Q_\star'\)</span> for the given binary are considered converged (see article text).</p> <pre><code>&lt;CLUSTER&gt;_&lt;BINARY ID&gt;_cdfstd_period.mrt</code></pre> <p>The standard deviation of the <span class="math-tex">\(CDF(\log_{10}Q_\star')\)</span> for the given binary as a function of tidal period for each of the quantiles. The maximum likelihood value is the target percentile, i.e. one of: 2.3%, 15.9%, 84.1%, and 97.7%</p>

opencc-by-4.0May 2022View details →
zenodo40/100

◂Fig. 6 A molecular phylogeny of 56 systematically representative Peridiniaceae, including 42 accessions assignable to P. cinctum from various geographic regions. Maximum likelihood tree (– ln = 21,884.93), as inferred from a rRNA nucleotide alignment (1137 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (CZE Czech Republic, E East, GER Germany, HET Heterocapsaceae, N North, PPE Protoperidiniaceae, POL Poland, rbn ribotype n, S South, SWE Sweden, UKR Ukraine, W West) in Bumps on the back: An unusual morphology in phylogenetically distinct Peridinium aff. cinctum (= Peridinium tuberosum; Peridiniales, Dinophyceae)

◂Fig. 6 A molecular phylogeny of 56 systematically representative Peridiniaceae, including 42 accessions assignable to P. cinctum from various geographic regions. Maximum likelihood tree (– ln = 21,884.93), as inferred from a rRNA nucleotide alignment (1137 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (CZE Czech Republic, E East, GER Germany, HET Heterocapsaceae, N North, PPE Protoperidiniaceae, POL Poland, rbn ribotype n, S South, SWE Sweden, UKR Ukraine, W West)

opencc-by-4.0Jan 2024View details →
zenodo40/100

◂Fig. 4 A molecular tree of 51 systematically representative Peridiniaceae, including all 28 accessions assignable to P. volzii. Maximum Likelihood tree (–ln = 22,017.62), as inferred from a rRNA nucleotide alignment (1,129 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (abbreviations: HET, Heterocapsaceae; PPE, Protoperidiniaceae) in Morphological and molecular variability of Peridinium volzii Lemmerm. (Peridiniaceae, Dinophyceae) and its relevance for infraspecific taxonomy

◂Fig. 4 A molecular tree of 51 systematically representative Peridiniaceae, including all 28 accessions assignable to P. volzii. Maximum Likelihood tree (–ln = 22,017.62), as inferred from a rRNA nucleotide alignment (1,129 parsimony-informative sites) and with strain number information. Numbers on branches are ML bootstrap (above) and Bayesian support values (below) for the clusters (asterisks indicate maximal support values, values under 50 and 0.90, respectively, are not shown). Clades are indicated (abbreviations: HET, Heterocapsaceae; PPE, Protoperidiniaceae)

opencc-by-4.0Oct 2021View details →
zenodo36/100

Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics: 35 Binding Site Data

<p>This is the original data for the manuscript &quot;Diffraction-Limited Molecular Cluster Quantification with Bayesian Nonparametrics&quot; by J Bryan IV, I Sgouralis, and S Presse. This repository contains movies of DNA origami with 35 binding sites.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

A Clustering Approach to Improve IntraVoxel Incoherent Motion Maps from DW-MRI using Conditional Auto-Regressive Bayesian Model

<p>Simulated data generated and used in the paper &quot;A Clustering Approach to Improve IntraVoxel Incoherent Motion Maps from DW-MRI using Conditional Auto-Regressive Bayesian Model&quot; are here available.</p> <p>Results generated from both simulated and clinical datasets are also available on the excel tables.</p>

opencc-by-4.0Jan 2022View details →
dryad32/100

Data from: Bayesian clustering analyses for genetic assignment and study of hybridization in oaks: effects of asymmetric phylogenies and asymmetric sampling schemes

Bayesian clustering methods have been widely used for studying species delimitation and genetic introgression. In order to test the effect of phylogenetic relationships and sampling scheme on the inferred clustering solution and on the performance of Bayesian clustering analysis, I simulated genotypes of the interfertile oak species Quercus robur, Quercus petraea, and Quercus pubescens and I run analyses using two popular software programs, STRUCTURE and BAPS. First, based on purebred simulations, I compared clustering solutions resulting from different sample size configurations. While clustering solution generally reflected the taxonomic relationships when equal samples of each species were included, spurious partition was inferred by STRUCTURE when some species were represented by larger and others by smaller samples. In very unbalanced configurations, STRUCTURE failed to identify the three species, even if three subpopulations were assumed. By contrast, BAPS could properly identify the three species under any sampling scheme. Second, based on simulations of purebreds and hybrids, I tested the performance of individual assignments with variable number of loci. This analysis showed that STRUCTURE can detect introgressed individuals more efficiently than BAPS. However, BAPS could assign purebreds more efficiently with a lower number of loci. Method performance also depended on phylogenetic relationships. In the case of Q. petraea, Q. pubescens, and their hybrids, method performance was lower due to their phylogenetic affinity. Inclusion of three instead of two species into the analysis led to reduction of performance, and to misclassification of hybrids, which often reflected the phylogenetic affinity between Q. petraea and Q. pubescens.

opencc-zeroDec 2012View details →
zenodo32/100

BASCULE: Bayesian inference and clustering of mutational signatures leveraging biological priors

<p>In the preprint available at https://doi.org/10.1101/2024.09.16.613266 we present BASCULE, a new method to perform Bayesian signatures deconvolution and to cluster patients from the inferred exposures. The method can deconvolve any kind of mutational signature types (SBS, DBS, ID, etc.) including as input a reference catalogue of known signatures (i.e., COSMIC), and cluster the samples joinltly from the exposures of all signature types. BASCULE is available as an R package (https://github.com/caravagnalab/bascule.git). Here we release the data and code to reproduce the analysis on synthetic and real datasets presented in the preprint, in the "synthetic_data_validation.zip" and "real_data_validation.zip", respectively.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

FIGURE 53. Bayesian clustering for K in A taxonomic revision of the Palaearctic species of the ant genus Tapinoma Mayr 1861 (Hymenoptera: Formicidae)

FIGURE 53. Bayesian clustering for K=2 of 15 microsatellite loci of Tapinoma madeirense (orange) and T. subboreale (blue) in southern France with hybridization in a 100 km wide zone along the Rhone river from about Nîmes to Montélimar.

opennotspecifiedApr 2024View details →
zenodo32/100

FIGURE 51. Bayesian clustering for K in A taxonomic revision of the Palaearctic species of the ant genus Tapinoma Mayr 1861 (Hymenoptera: Formicidae)

FIGURE 51. Bayesian clustering for K=2 of 15 microsatellite loci of Tapinoma madeirense (orange) and T. subboreale (blue).

opennotspecifiedApr 2024View details →
zenodo32/100

Exploring the conformers of an organic adsorbate on a metal cluster with Bayesian optimization

<p>[1] Lincan Fang, Xiaomi Guo, Milica Todorovic, Patrick Rinke, Xi Chen. Exploring the conformers of an organic adsorbate on a metal cluster with Bayesian optimization</p> <p>This dataset shows 1000 sampling points in the configuration space of the cysteine ligand binding with the S site of Au25S18 cluster. For Au25S18 cluster, there are two inequivalent S sites (Au-S-Au-S-Au-S-Au arrow top and side). Here, the binding S site is on &quot;arrow side&quot;, we named it &quot;system B&quot;. The sampling method is the Bayesian Optimization Structure Search (BOSS), each sampling point consists of structure features and DFT structure energy. The structure features are five dihedral angles of cysteine ligand, d1 (Au-S-C1-C2), d2 (S-C1-C2-N), d3 (C1-C2-N-H), d4 (C1-C2-C3-O1), and d5 (C2-C3-O1-H). The energy was calculated by FHI-aims with PBE functional, tier 2 setting with many-body dispersion corrections. Due to the confined configuration space of this system, we use the energy transformation method (see in manuscript) to tackle one sampling structure that cannot be simulated by DFT or has a huge high DFT energy. For additional details on BOSS search please refer to [1].</p> <p>The stable local minimum structures from BOSS can be found in NOMAD:&nbsp;https://dx.doi.org/10.17172/NOMAD/2022.08.20-1</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
dryad32/100

Data from: Evaluating the ability of Bayesian clustering methods to detect hybridization and introgression using an empirical red wolf dataset

Open the record for dataset details and reuse information.

publicOct 2012View details →
dryad32/100

Data from: Bayesian clustering analyses for genetic assignment and study of hybridization in oaks: effects of asymmetric phylogenies and asymmetric sampling schemes

Open the record for dataset details and reuse information.

publicNov 2014View details →
dryad28/100

Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis

The inference of population genetic structures is essential in many research areas in population genetics, conservation biology and evolutionary biology. Recently, unsupervised Bayesian clustering algorithms have been developed to detect a hidden population structure from genotypic data, assuming among others that individuals taken from the population are unrelated. Because of this hypothesis, markers in a sample taken from a subpopulation can be considered to be in Hardy-Weinberg and linkage equilibrium. However, close relatives might be sampled from the same subpopulation, and consequently, might cause Hardy-Weinberg and linkage disequilibrium and thus bias a population genetic structure analysis. In this study, we used simulated and real data to investigate the impact of close relatives in a sample on Bayesian population structure analysis. We also showed that, when close relatives were identified by a pedigree reconstruction approach and removed, the accuracy of a population genetic structure analysis can be greatly improved. The results indicate that unsupervised Bayesian clustering algorithms cannot be used blindly to detect genetic structure in a sample with closely related individuals. Rather, when closely related individuals are suspected to be frequent in a sample, these individuals should be first identified and removed before conducting a population structure analysis.

opencc-zeroDec 2011View details →
dryad28/100

The impact of estimator choice: Disagreement in clustering solutions across K estimators for Bayesian analysis of population genetic structure across a wide range of empirical datasets

<p class="CxSpFirst">The software program STRUCTURE is one of the most cited tools for determining population structure. To infer the optimal number of clusters from STRUCTURE output, the Δ<i>K</i> method is often applied. However, a recent study relying on simulated microsatellite data suggested that this method has a downward bias in its estimation of <i>K</i> and is sensitive to uneven sampling. If this finding holds for empirical datasets, conclusions about the scale of gene flow may have to be revised for a large number of studies. To determine the impact of method choice, we applied recently described estimators of <i>K</i> to re-estimate genetic structure in 41 empirical microsatellite datasets; 15 from a broad range of taxa and 26 focused on a diverse phylogenetic group, coral. We compared alternative estimates of <i>K</i> (Puechmaille statistics) with traditional (Δ<i>K</i> and posterior probability) estimates and found widespread disagreement of estimators across datasets. Thus, one estimator alone is insufficient for determining the optimal number of clusters regardless of study organism or evenness of sampling scheme. Subsequent analysis of molecular variance (AMOVA) between clustering solutions did not necessarily clarify which solution was best. To better infer population structure, we suggest a combination of visual inspection of STRUCTURE plots and calculation of the alternative estimators at various thresholds in addition to Δ<i>K</i>. Differences between estimators could reveal patterns with important biological implications, such as the potential for more population structure than previously estimated, as was the case for many studies reanalyzed here.</p>

opencc-zeroOct 2021View details →
dryad28/100

Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis

Open the record for dataset details and reuse information.

publicJun 2012View details →
dryad28/100

The impact of estimator choice: Disagreement in clustering solutions across K estimators for Bayesian analysis of population genetic structure across a wide range of empirical datasets

Open the record for dataset details and reuse information.

publicOct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record