Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,326
datasets available to search
ShareScore release 0.9.0
Dataset results
2,326 results for “clusters”
Data from: Hidden Markov models reveal tactical adjustment of temporally-clustered courtship displays in response to the behaviors of a robotic female
Open the record for dataset details and reuse information.
Data from: The palaeobiological significance of clustering in acritarchs: a case study from the early Cambrian of North Greenland
Open the record for dataset details and reuse information.
Phylogenomics of the psychoactive mushroom genus Psilocybe and evolution of the psilocybin biosynthetic gene cluster
Open the record for dataset details and reuse information.
Ancient gene clusters govern the initiation of monoterpenoid indole alkaloid biosynthesis and C3 stereochemistry inversion
Open the record for dataset details and reuse information.
Data from: Joint spatial modeling of cluster size and density for a heavily hunted primate persisting in a heterogeneous landscape
Open the record for dataset details and reuse information.
Regional E-Atlas of the Greater Phoenix Region: high-technology employment clusters, 2000.
These data represent high-technology employment clusters across central Arizona-Phoenix. These data are presented by industry: aerospace, bio-industry, information, software for the year 2000.
ESFRI thematic cluster view on EOSC
<p>The illustration visualizes the view of the five ESFRI cluster projects EOSC-Life, ENVRI-FAIR, SSHOC, PANOSC and ESCAPE on the emerging European Open Science Cloud. Particular focus is put on the three building layers of consolidated e-infrastructures, open science, and scientific communities' content and users which the ESFRI cluster projects bring to the EOSC. The three layers identified by the cluster projects provide an "alternative" interpretation of the acronym EOSC.</p>
The Qualitative Analysis of Repertory Grid Data: Interpretive Clustering (datasets only)
<p>Datasets accompanying the publication "The Qualitative Analysis of Repertory Grid Data: Interpretive Clustering" by </p> <p>Burr, King, and Heckmann.</p>
K-Means Clustering Peraturan Kementerian
<p>Analisis Peraturan Kementerian menggunakan metode K-Means Clustering</p>
Pairwise distance demarcation of species in the family Coronaviridae. a, Diagonal matrix of PPDs of 2,505 viruses clustered according to 49 coronavirus species, 39 established and 10 pending or tentative, and ordered from the most to least populous species, from left to right; green and white, PPDs smaller and larger than the inter-species threshold, respectively. Areas of the green squares along the diagonal are proportional to the virus sampling of the respective species, and virus prototypes of the five most sampled species are specified to the left; asterisks indicate species that include viruses whose intra-species PPDs crossed the inter-species threshold (threshold 'violators'). b, Maximal intra-species PPDs (x axis, linear scale) plotted against virus sampling (y axis, log scale) for 49 species (green dots) of the Coronaviridae. Indicated are the acronyms of virus prototypes of the seven most sampled species. Green and blue plot sections represent intra-species and intra-subgenera PPD ranges. The vertical black line indicates the inter-species threshold. c, Shown are the PDs of non-identical residues (y axis) for four viruses representing three major phylogenetic lineages (clades) of the species Severe acute respiratorysyndrome-related coronavirus (panel b) and all pairs of the 256 viruses of this species ('all pairs'). The PD values were derived from pairwise distances in the MSA that were calculated using an identity matrix. Panels a and b were adopted from the DEmARC v.1.4 output. in The species Severe acute respiratory syndromerelated coronavirus: classifying 2019-nCoV and naming it SARS-CoV-2
Pairwise distance demarcation of species in the family Coronaviridae. a, Diagonal matrix of PPDs of 2,505 viruses clustered according to 49 coronavirus species, 39 established and 10 pending or tentative, and ordered from the most to least populous species, from left to right; green and white, PPDs smaller and larger than the inter-species threshold, respectively. Areas of the green squares along the diagonal are proportional to the virus sampling of the respective species, and virus prototypes of the five most sampled species are specified to the left; asterisks indicate species that include viruses whose intra-species PPDs crossed the inter-species threshold (threshold 'violators'). b, Maximal intra-species PPDs (x axis, linear scale) plotted against virus sampling (y axis, log scale) for 49 species (green dots) of the Coronaviridae. Indicated are the acronyms of virus prototypes of the seven most sampled species. Green and blue plot sections represent intra-species and intra-subgenera PPD ranges. The vertical black line indicates the inter-species threshold. c, Shown are the PDs of non-identical residues (y axis) for four viruses representing three major phylogenetic lineages (clades) of the species Severe acute respiratorysyndrome-related coronavirus (panel b) and all pairs of the 256 viruses of this species ('all pairs'). The PD values were derived from pairwise distances in the MSA that were calculated using an identity matrix. Panels a and b were adopted from the DEmARC v.1.4 output.
PAN17 Author Identification: Clustering
<p>We provide a collection of (up to 50) short documents (paragraphs extracted from larger documents), identify authorship links and groups of documents by the same author. All documents are single-authored, in the same language, and belong to the same genre. However, the topic or text-length of documents may vary. The number of distinct authors whose documents are included in the collection is not given.</p> <p>More information: <a href="https://pan.webis.de/clef17/pan17-web/author-clustering.html">Link</a></p>
Supplementary data for Cariou et al (2020, Molecular Ecology Resources, "How consistent is RAD-seq divergence with DNA-barcode based clustering in insects?")
<p>This dataset accompanies a paper by Cariou et al, to be published in Molecular Ecology Resources, where we assessed in 92 insect species if the genetic clustering of specimens into species like units, on the basis of mitochondrial DNA, was consistent with genome wide divergence, as estimated by RAD-seq data. The present repository includes: (1) a detailed description of the bioinformatic analysis indicating which programs were used, together with parameter values, (2) the raw RAD-seq data following demultiplexing, (3) the consensus sequences of all RAD loci for all specimens, and (4) large tables indicating genetic distances at all RAD loci for all species.</p>
FIGURE5. Maximum likelihood tree based on the Kimura 2-parameter model of the COI sequences from the Siphamia species with P. kauderni as the outgroup. Tree shown here has the highest log likelihood following 10 000 replications. The percentage of trees in which the associated taxa clustered together is shown next to the branches, branch lengths are measured in the number of substitutions per site and all positions containing gaps and missing data have been eliminated. in Redescription and distributional range extension of the Speckled Siphonfish, Siphamia guttulata (Pisces: Apogonidae)
FIGURE5. Maximum likelihood tree based on the Kimura 2-parameter model of the COI sequences from the Siphamia species with P. kauderni as the outgroup. Tree shown here has the highest log likelihood following 10 000 replications. The percentage of trees in which the associated taxa clustered together is shown next to the branches, branch lengths are measured in the number of substitutions per site and all positions containing gaps and missing data have been eliminated.
WiDiv_PAM_Clustering_Integrated_Phenotypes
<p>Root phenotype and performance data for the Wisconsin Diversity Panel collected from the field under well-watered and water stress conditions in Willcox, AZ, in 2016. Images of root architecture and anatomy were analyzed for each genotype, two replicates per treatment. Architectural data was collected using DIRT. Anatomical data was collected with RootScan2 and MIPAR. Also included is the R script to conduct a PAM clustering analysis to identify clusters of root phenotypes related to performance.</p>
Clustering datasets
<p>Iris, Moon, and Circles datasets for Galaxy clustering tutorial </p>
Interdependency of port clusters during regional disasters
<p>Ports play a vital role in the economy of nations and provide a critical link in the supply chain. Ports form the gateway by which essential goods are received within large geographic regions. Because of their function, ports are exposed to a substantial risk of flooding, storm events, sea-level rise, and climate change. The resiliency of ports is essential for the economy, the people, and national readiness. The contribution of this research work is in providing a methodology to quantify port resiliency that is applicable at the individual port level and regionally. The research approach first defines a quantifiable measure of systematic resiliency. Then applies this measure to quantify the resiliency of six ports located in the Southeast US impacted by Hurricane Matthew (2016). The details are presented in the final report. Also, the data set used for the report is included.</p>
Original data for: Deracemization of Au38 Clusters: Breaking the Equal Status and Dynamic Inversion of Enantiomers
<p>Original data for Figures 2, 6, S1, S2, S3, S4, S5, S6, S7, S8, S9</p>
Strong chemical tagging with APOGEE: 21 candidate star clusters that have dissolved across the Milky Way disc
<p>Two files containing the chemically tagged abundances derived by astroNN from the Apache Point Observatory Galactic Evolution Experiment. The chemical tagging procedure is described in <a href="https://arxiv.org/abs/2004.04263">https://arxiv.org/abs/2004.04263</a>. DBSCAN_labels_APOGEE_DR16.npy contains only members of groups with more than 15 members and silhouette coefficients greater than 0. DBSCAN_labels_APOGEE_DR16_all.npy contains labels for all stars in the quality controlled data set. The data are packaged as a numpy structured array. Most of the columns are described at <a href="https://data.sdss.org/datamodel/files/APOGEE_ASTRONN/apogee_astronn.html">https://data.sdss.org/datamodel/files/APOGEE_ASTRONN/apogee_astronn.html</a>. There are two additional columns: CLUSTER_ID, and SILHOUETTE_COEFF, which are respectively the label and silhouette coefficient for each star.</p>
Asynchronous event-based clustering and tracking for intrusion monitoring in UAS
<p>This dataset describes a collection of rosbag files for event-based intruder monitoring using UAS. A DAVIS346 camera was mounted over a DJI Flamewheel F550 Drone, and an onboard computer recorded the sensor information from the event camera. Each dataset includes events, frames, and IMU measurements. The monitoring scenes were recorded outdoors at the School of Engineering of the University of Seville. In each dataset, an intruder moves and hides from the field of view of the camera simulating a scape-intrusion situation. A total of four monitoring setting were recorded:</p> <p><strong>Daylight monitoring:</strong> A daylight scene for intruder monitoring. An intruder runs and hides behind the objects of the scene to evade the camera field of view.<br> <br> <strong>Night light monitoring:</strong> A monitoring scene during the night without the presence of any artificial light. The low light condition increases the difficulty of monitoring task due to the increment of noisy events.<br> <br> <strong>Multi-target:</strong> An experiment with a suspect and a chaser drone moving in the monitoring area. The drone follows the suspect by simulating a pursuit operation.<br> <br> <strong>Monitoring under illumination changes:</strong> A night scene where the lighting conditions changes by the movement of artificial lights in the scene.</p>
Neisseria gonorrhoeae clustering to reveal major European WGS-based genogroups in association with antimicrobial resistance (cgMLST and MScgMLST schemas, allelic profile matrices and GrapeTree input file)
<p>This dataset refers to the gene-by-gene analysis of 3791 <em>Neisseria gonorrhoeae</em> genomes from 21 European countries and includes the used cgMLST and MScgMLST loci schemas prepared for the chewBBACA core suite, as well as the associated allelic profile matrices for all genomes. Additionally a <em>.json</em> file is made available for direct input in the GrapTree vizualization software for data/metadata exploration. </p> <p>All novel raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB36482). Additional raw sequence read data used were retrieved from the following ENA BioProjects: PRJEB14933; PRJEB2124; PRJEB23008; PRJEB26560; PRJEB9227; PRJNA275092; PRJNA348107; PRJNA473385; PRJNA315363. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.