Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,326
datasets available to search
ShareScore release 0.9.0
Dataset results
2,326 results for “clusters”
Spatial clustering of trumpetfish shadowing behaviour in the Caribbean Sea revealed by citizen science
<p>The West Atlantic trumpetfish (Aulostomus maculatus) performs an unusual hunting strategy, termed shadowing, whereby a trumpetfish swims closely behind or next to another 'host' species to facilitate the capture of prey. Despite trumpetfish being observed throughout the Caribbean, observations of this behaviour appear to be concentrated to a handful of localities. Here we assess the degree of geographical clustering of shadowing behaviour throughout the Caribbean Sea, and identify ecological features associated with the likelihood of its occurrence. To do this, we used a citizen science approach by creating and distributing an online survey to target frequent divers across this region. While the vast majority of participants observed trumpetfish on nearly every dive across the Caribbean, using random labelling spatial analyses, we found the frequency of shadowing behaviour was geographically clustered; participants that were within ~ 120 km of each other reported observations of shadowing that were more similar than would be expected by chance. Our survey also highlighted that trumpetfish were more likely to be observed shadowing than observed alone in a particular habitat type, and with particular host species, suggesting potential ecological factors that could drive the uneven distribution of this behaviour. Our results demonstrate that this behavioural hunting strategy is spatially clustered and, more generally, highlight the power of using citizen science to investigate variation in animal behaviour over thousands of square kilometres.</p>
The R136 star cluster dissected with Hubble Space Telescope/STIS. III. The most massive stars and their clumped winds
<p>Supplemental material to 'The R136 star cluster dissected with Hubble Space Telescope/STIS. III. The most massive stars and their clumped winds'. This includes, for each star in the sample: a plot of the observed spectrum overlaid with best fitting models and fitness diagrams for all free parameters. Furthermore we include data (the normalised spectra used in the analysis, and best fitting models) and a basic reproduction package. </p>
Data Supporting "Mesoscale Convective Clustering Enhances Tropical Precipitation"
<p>These data are in support of "Mesoscale Convective Clustering Enhances Tropical Precipitation" by P. Angulo-Umana and D. Kim. </p> <p>In the tropics, extreme precipitation events are often caused by mesoscale systems of organized, spatially clustered deep cumulonimbi, posing a substantial risk to life and property. While the clustering of convective clouds has been thought to strengthen precipitation intensity, no quantitative estimates of this hypothesized enhancement exist. In this study, after isolating the effects of mesoscale convective clustering on precipitation, we find that strongly clustered convection precipitates more intensely than weakly clustered convection. We further show that this enhancement is primarily attributable to an increase in convective precipitation intensity when the environment is less than 70% saturated, with increases in stratiform cloud cover being of equal or greater importance when the environment is closer to saturation. Our results suggest that a correct representation of mesoscale organized convective systems in numerical weather and climate models is needed for accurate predictions of extreme precipitation events.</p>
Data from: Stepwise Threshold Clustering: a new method for genotyping MHC loci using next-generation sequencing technology
Genes of the vertebrate major histocompatibility complex (MHC) are of great interest to biologists because of their important role in immunity and disease, and their extremely high levels of genetic diversity. Next generation sequencing (NGS) technologies are quickly becoming the method of choice for high-throughput genotyping of multi-locus templates like MHC in non-model organisms. Previous approaches to genotyping MHC genes using NGS technologies suffer from two problems: 1) a "gray zone" where low frequency alleles and high frequency artifacts can be difficult to disentangle and 2) a similar sequence problem, where very similar alleles can be difficult to distinguish as two distinct alleles. Here were present a new method for genotyping MHC loci – Stepwise Threshold Clustering (STC) – that addresses these problems by taking full advantage of the increase in sequence data provided by NGS technologies. Unlike previous approaches for genotyping MHC with NGS data that attempt to classify individual sequences as alleles or artifacts, STC uses a quasi-Dirichlet clustering algorithm to cluster similar sequences at increasing levels of sequence similarity. By applying frequency and similarity based criteria to clusters rather than individual sequences, STC is able to successfully identify clusters of sequences that correspond to individual or similar alleles present in the genomes of individual samples. Furthermore, STC does not require duplicate runs of all samples, increasing the number of samples that can be genotyped in a given project. We show how the STC method works using a single sample library. We then apply STC to 295 threespine stickleback (Gasterosteus aculeatus) samples from four populations and show that neighboring populations differ significantly in MHC allele pools. We show that STC is a reliable, accurate, efficient, and flexible method for genotyping MHC that will be of use to biologists interested in a variety of downstream applications.
Distribution. Found in three main clusters in C Cameroon and SW Central African Republic (both are known only from single records), W Kenya, and S & E South Africa and Swaziland; range is probably much more extensive, although it has not been collected readily outside its South African and Kenyan distribution. in Soricidae
Distribution. Found in three main clusters in C Cameroon and SW Central African Republic (both are known only from single records), W Kenya, and S & E South Africa and Swaziland; range is probably much more extensive, although it has not been collected readily outside its South African and Kenyan distribution.
Exploring the conformers of an organic adsorbate on a metal cluster with Bayesian optimization
<p>[1] Lincan Fang, Xiaomi Guo, Milica Todorovic, Patrick Rinke, Xi Chen. Exploring the conformers of an organic adsorbate on a metal cluster with Bayesian optimization</p> <p>This dataset shows 1000 sampling points in the configuration space of the cysteine ligand binding with the S site of Au25S18 cluster. For Au25S18 cluster, there are two inequivalent S sites (Au-S-Au-S-Au-S-Au arrow top and side). Here, the binding S site is on "arrow side", we named it "system B". The sampling method is the Bayesian Optimization Structure Search (BOSS), each sampling point consists of structure features and DFT structure energy. The structure features are five dihedral angles of cysteine ligand, d1 (Au-S-C1-C2), d2 (S-C1-C2-N), d3 (C1-C2-N-H), d4 (C1-C2-C3-O1), and d5 (C2-C3-O1-H). The energy was calculated by FHI-aims with PBE functional, tier 2 setting with many-body dispersion corrections. Due to the confined configuration space of this system, we use the energy transformation method (see in manuscript) to tackle one sampling structure that cannot be simulated by DFT or has a huge high DFT energy. For additional details on BOSS search please refer to [1].</p> <p>The stable local minimum structures from BOSS can be found in NOMAD: https://dx.doi.org/10.17172/NOMAD/2022.08.20-1</p> <p> </p>
Genome sequences of Rhizopogon roseolus, Mariannaea elegans, Myrothecium verrucaria, and Sphaerostilbella broomeana and the identification of biosynthetic gene clusters for fungal peptide natural products
<p>Data to accompany the paper, detailed in manifest.txt</p>
Double-lined spectroscopic binaries in the open cluster M 11 (NGC 6705)
<p>Figures with spectral fits for 265 Gaia-ESO spectra for open cluster M 11 (NGC 6705), analysed by single-star and binary spectroscopic models. </p> <p>Kovalev M., Straumit I., 2022, MNRAS, 510, 1515</p> <p>also used in arxiv 2207:06996</p>
Deducing Subnanometer Cluster Size and Shape Distributions of Heterogeneous Supported Catalysts
<p>Dataset for "Deducing Subnanometer Cluster Size and Shape Distributions of Heterogeneous Supported Catalysts". </p>
Clustering Data for Passive Tracers in Random Velocity Fields
<p>Contains the post-processed model output. All detected clusters from each individual model run used to compile statistics are included as well as python scripts to manipulate cluster objects and create figures.</p>
Supplementary material 1 from: Wang J-h, Zheng X-d (2017) Comparison of the genetic relationship between nine Cephalopod species based on cluster analysis of karyotype evolutionary distance. Comparative Cytogenetics 11(3): 477-494. https://doi.org/10.3897/compcytogen.v11i3.12752
Chromosome relative length, supplemental formulae : Explanation note: Chromosome relative length, supplemental formulae and all of the original images are made available under the online digital repository Figshare, and it is free to access, in adherence to the principle of open data, more details in https://figshare.com/s/8d21a0db9ffe1f17d279
Precomputed OOPS dataset for Actinobacterial clusters output from antiSMASH
<p>This dataset is the precomputed actinobacterial clusters taken from the antiSMASH database, ready to be loaded into OOPS.</p>
Experimental microwave spectroscopic and quantum-chemical investigation of the HCl(H2O)n (n=1-5,7) molecular clusters
<p>In this dataset, experimental microwave rotational spectroscopy data and corresponding theoretical calculations for the HCl(H2O)n (n=1-5,7) clusters are stored. </p>
Intermolecular interactions probed by rotational dynamics in gas-phase clusters
<p>all the raw data for the main figures of our literature "Intermolecular interactions probed by rotational dynamics in gas-phase clusters"</p>
Experimental data: "Dynamical clustering and wetting phenomena in inertial active matter".
<p>In each zip folder, we report experimental data for shaker frequency f=90 Hz, 120 Hz, 150 Hz at different packing fraction, corresponding to different number of particles N=30, 60, 90, 120, 150, 180, 210, 240</p>
The developmental and evolutionary characteristics of transcription factor binding site clustered regions based on an explainable machine learning model
<p>## Identification of transcription factor binding sites clustered regions</p> <p>First, the TFBSs were identified from ATAC-seq peaks by FIMO. The position-specific weight matrices (PWMs) of transcription factors were downloaded from CIS-BP databases. The genomic sequences under the open chromatin regions were used as inputs for FIMO with a custom library of all motifs for each species to scan for motif instances at a p-value threshold of 1e-5. </p> <p>Then, an established method was used to identify TFCRs by performing the Gaussian kernel density estimations across the genome (with a bandwidth of 300bp centered on each TFBS). Each peak in density profile was considered a TFCR. To determine the complexity of each TFCR, the Gaussian kernelized distances from each peak that contributed at least 0.1 to its strength were determined. The complexity of each TFCR was determined by the quantity and proximity of the contributing TFBS. We combined motif instances based on the TF family information from CIS-BP to calculate the complexity of TFCR. The window for each TFCR was determined by finding the maximum distance (in bp) from the TFCR to a contributing TF and then adding 150 bp (one-half of the bandwidth). Each window was centered on the TFCR. The identified TFCR was grouped into 10 groups based on their complexity from low to high. </p> <p>usage: <br>indir="Human_fimo" # the directory where you put the output files of FIMO <br>motifMap="Homo_sapiens_2020_0920/TF_Information_all_motifs_plus.txt" # the mapping relationship of TF and its TF family from CIS-BP <br>cd Codes/TFCR_embryo <br>perl d-motif_combine.pl $indir TFfamily $motifMap <br>perl e-tfpos_combine.pl TFfamily <br>perl f1-tf_bed-new-c.pl TFfamily <br>perl 0-merge-TFCR.pl $indir TFfamily </p>
Data availability: Clustered and rotating designs as a strategy to obtain precise detection rates in camera trapping studies
<p>Manuscript data "Clustered and rotating designs as a strategy to obtain precise detection rates in camera trapping studies" published in Journal of Applied Ecology. R code to replicate the simulations can be found in the supplementary materials of the manuscript.</p>
Instances of the problem of Designing a Multi-sink Clustered Wireless Sensor Network.
Open the record for dataset details and reuse information.
A Layer-Elastic Scheduling System for DLT in GPU Clusters
<p>This artifact contains the simulator and K8S implementation for a layer-elastic scheduling system. </p>
Core Binding Energy Calculations: A Scalable Approach with the Quantum Embedding Based Equation-of-Motion Coupled-Cluster Method
<p>This data includes the HF-optimized orbitals, coupled cluster amplitudes (T1, T2), and EOM-CCSD left and right eigenvectors at the CC-PCVDZ basis set. It can be used to reproduce the data for "Core Binding Energy Calculations: A Scalable Approach with the Quantum Embedding Based Equation-of-Motion Coupled-Cluster Method."</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.