Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,326

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,326 results for “clusters”

Learn how ShareScore rates datasets ↗
zenodo36/100

Cataloging Distant Galactic Open Clusters: Identification of 739 New Star Clusters Beyond 5 kpc Utilizing GAIA DR3 Data

<p><span>The figures of the 739 open clusters that report in our paper (Cataloging Distant Galactic Open Clusters: Identification of &nbsp;739 New Star Clusters Beyond 5 kpc Utilizing GAIA DR3 Data)</span></p> <p><span>&nbsp;complete King's model<span>&nbsp; </span>profile fitting, sky charts and 5-panels ( spatial distribution, proper-motion distribution,</span></p> <p><span><span>parallax statistics, parallax distribution, and CMD)</span> .</span></p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Code and data for spatial and temporal magnitude clustering analysis

<p>Code used for performing spatial and temporal&nbsp;seismic magnitude clustering analysis.&nbsp; Includes documentation (README.txt) with steps on how to implement the code. The public datasets used for this study can be accessed at the following locations:&nbsp;</p> <ul> <li><strong>Southern California Catalog:&nbsp;</strong> <ul> <li>SCEDC (2013): Southern California Earthquake Center.<br> Caltech.Dataset. doi:<a href="https://dx.doi.org/10.7909/C3WD3xH1">10.7909/C3WD3xH1</a></li> </ul> </li> <li><strong>Northern California Catalog:</strong> <ul> <li>NCEDC (2014), Northern California Earthquake Data Center. UC Berkeley Seismological Laboratory. Dataset. doi:10.7932/NCEDC.</li> </ul> </li> <li><strong>Mixed-mode Laboratory Catalog:</strong> <ul> <li>Lin, Qing, et al. &quot;Opening and mixed mode fracture processes in a quasi-brittle material via digital imaging.&quot;&nbsp;<em>Engineering Fracture Mechanics</em>&nbsp;131 (2014): 176-193.</li> </ul> </li> <li><strong>ETAS Code:</strong> <ul> <li>Leila Mizrahi, Shyam Nandan, Stefan Wiemer 2021;<br> Embracing Data Incompleteness for Better Earthquake Forecasting. (Section 3.1)<br> <em>Journal of Geophysical Research: Solid Earth</em>; doi:&nbsp;<a href="https://doi.org/10.1029/2021JB022379">https://doi.org/10.1029/2021JB022379</a></li> </ul> </li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Clust&See3.0 : clustering, module exploration and annotation.

<p><span>Clust&amp;See</span><span>3</span><span>.0 is the novel version of a </span><span>Cytoscape</span><span> app </span><span>that has been developed to </span><span>identify, visualize and manipulate network clusters and modules, newly enriched with functionalities allowing </span><span>custom </span><span>annotation</span><span>s</span><span> of nodes and computation of their statistical enrichment</span><span>s</span><span>. As the wealth of multi-omics data is growing,</span><span> such functionalities are highly valuable for a better understanding of biological module composition</span><span>.</span></p> <p>&nbsp;</p> <ul> <li><span>go_annot.txt : The protein annotation file extracted from Gene Ontology Biological Process database.</span></li> <li><span>HuRI_CC.txt : The network file containing the largest connect component of&nbsp; <span>the&nbsp;</span>human reference interactome network [1]</span></li> </ul> <p><span>[1] <span>Luck K, Kim D-K, Lambourne L, Spirohn K, Begg BE, Bian W, et al. A reference map of the human binary protein interactome. Nature. 2020;580:402&ndash;8. </span></span></p>

opencc-by-4.0Jun 2024View details →
dryad36/100

Deciphering the explanatory potential of blood pressure variables on post-operative length of stay through hierarchical clustering: A retrospective monocentric study

<p><em>Objective:</em> Mean arterial pressure is widely used as the variable to monitor during anesthesia. But there are many other variables proposed to define intraoperative arterial hypotension. The goal of the present study was to search arterial pressure variables linked with prolonged postoperative length of stay (pLOS).</p> <p><em>Design: </em>Retrospective cohort study of adult patients having received general  for a scheduled non cardiac surgical procedure between 15<sup>th</sup> July 2017 and 31st December 2019.</p> <p><em>Methods:</em> pLOS was defined as a stay longer than the median (main outcome), adjusted for surgery type and duration. 330 arterial pressure variables were analyzed and organized through a clustering approach. An unsupervised hierarchical aggregation method for optimal cluster determination, employing Kendall's tau coefficients and a penalized Bayes information criterion was used. Variables were ranked using the absolute standardized mean distance (aSMD) to measure their effect on pLOS. Finally, after multivariate independence analysis, the number of variables was reduced to three.</p> <p><em>Results:</em> Our study examined 9,516 patients. When LOS is defined as strictly greater than the median, 34% of patients experienced pLOS. Key arterial pressure variables linked with this definition of pLOS included the difference between the highest and lowest pulse pressure values computed throughout the surgery (aSMD[95%CI] =0.39[0.31-0.40], p&lt;0.001), the accumulated time pulse pressure above 61mmHg (aSMD = 0.21[0.17-0.25], p&lt;0.001), and the lowest MAP during surgery (aSMD= 0.20[0.16-0.24], p&lt;0.001).</p> <p><em>Conclusions: </em>By applying a clustering approach, three arterial pressure variables were associated with pLOS. This scalable method can be applied to various dichotomized outcomes.</p>

opencc-zeroJul 2024View details →
zenodo36/100

Clustering Tasks and Decision Trees with Elegiac Poets

<p>The dataset contains files generated during a Natural Language Processing (NLP) and automatic text analysis task. Attached is a <strong>Jupyter notebook</strong> with the complete code, along with several <strong>Excel files (.xlsx)</strong> containing organized information. Additionally, there are three folders that include files generated during the Silhouette calculation, K-means clustering, and feature extraction using decision trees.</p> <p>The three folders are:<br>1. <strong>Silhouette Calculation:</strong> Contains PNG images of Silhouette plots for various analysis configurations.<br>2.<strong> K-means Clustering: </strong>Contains pickle (.pkl) files with features and labels for each combination of excluded author, n-gram type, n-gram range, and matrix type.<br>3. <strong>Feature Extraction: </strong>Contains CSV files with lists of documents by cluster and the most important features along with information gain and information gain ratio metrics.</p> <p>Other file formats included in the dataset are:<br>- CSV files containing Silhouette scores, optimal clustering results, cluster assignments, and optimal cluster assignments.<br>- PNG images of scatter plots colored by author and by cluster.<br>- Pickle files containing the top features extracted during the analysis.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Figure 3. Cluster Dendrogram indicating 11 in Composition and structure of plant communities in the Moist Temperate Forest Ecosystem of the Hindukush Mountains, Pakistan

Figure 3. Cluster Dendrogram indicating 11 plant association types in the Lalkoo Valley.

opencc-by-4.0Dec 2022View details →
zenodo36/100

Benchmark cancer datasets for Clustering algorithms for Omics-based Patient Stratification (COPS)

<p>This repository contains seven multi-omic cancer datasets including several cancer types (breast, kidney, lung, ovary, prostate, and thyroid cancers as well as low grade gliomas) that were used for benchmarking several multi-view clustering algorithms implemented by COPS (https://github.com/UEFBiomedicalInformaticsLab/COPS). The datasets were originally compiled from The Cancer Genoma Atlas (TCGA) and downloaded using the <em>curatedTCGAData</em> R-package. The datasets include copy-number variations, methylomics as well as mRNA and miRNA transcriptomics. The methylomics data was mapped to genes by averaging methylation level of probes associated with the promoter regions of genes. Similarly the miRNA transcriptomics data was mapped to genes by using known and predicted miRNA -&gt; gene interactions. Updated survival data was acquired from the Liu et al. 2018 paper.&nbsp;</p> <p>This repository also includes two sets of cancer associated pathway networks used by pathway-based multi-omic methods benchmarked in our study. NCI-PID pathways were downloaded using the <em>ndexr</em> R-package on December 22 2021. While KEGG pathways were downloaded using the <em>pathview</em> R-package on May 3 2022.&nbsp;</p> <p>More details on the processing can be found on the related publication.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Galaxy populations in the Hydra I cluster from the VEGAS survey III. The realm of low-surface brightness features and intra-cluster light

<p><span>This appendix provides the azimuthally-averaged surface bright</span><span>ness and colour profiles of the sample galaxies listed in Table 1 of the paper "Galaxy populations in the Hydra I cluster from the VEGAS survey III. The realm of low-surface brightness features and intra-cluster light", by <span>Marilena Spavone</span><span>,</span><span> Enrichetta Iodice</span><span>,</span><span> Felipe S. Lohmann</span><span>, Magda Arnaboldi</span><span>, Michael Hilker</span><span>, Antonio La&nbsp;</span><span>Marca</span><span>, Rosa Calvi</span><span>, Michele Cantiello</span><span>, Enrico M. Corsini</span><span>, Giuseppe D&rsquo;Ago</span><span>, Duncan A. Forbes</span><span>, Marco&nbsp;</span><span>Mirabile</span><span>, and Marina Rejkuba.</span><br></span></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Data file needed for star cluster setup in Phantom smoothed particle hydrodynamics and magnetohydrodynamics code

<p>** this file is automatically downloaded by Phantom when running the starcluster setup **</p> <p>This is a small ascii file containing positions and velocities of stars utilised in the "starcluster" configuration in the Phantom smoothed particle hydrodynamics and magnetohydrodynamics code (<a href="http://adsabs.harvard.edu/abs/2018PASA...35...31P">Price et al. 2018</a>). It is used to set up a collection of N-body particles.</p> <p>The star cluster setup (and the data file) were written by Yann Bernard as part of his PhD thesis at<strong> </strong>Universit&eacute; Grenoble Alpes. The datafile is published here so it can be used in the automated code testing via github actions.&nbsp;</p> <p>The columns are mass, position (x,y,z) and velocity (vx,vy,vz) for all of the stars in the simulation</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Original NGS dataset from publication "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center".

<p>Original hepatitis C virus hypervariable region 1 NGS sequences&nbsp;in fastq format from patients analyzed in the study&nbsp; &quot;Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center&quot;.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2018View details →
zenodo36/100

Influence of temperature on the molecular composition of ions and charged clusters during pure biogenic nucleation

<p>Data of Figures in manuscript:</p> <p>Influence of temperature on the molecular composition of ions and charged clusters during pure biogenic nucleation.</p> <p>Frege et al., Atmos. Chem. Phys. 18, 65&ndash;79, 2018</p> <p>https://doi.org/10.5194/acp-18-65-2018</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2018View details →
zenodo36/100

International Relations Articles Clustering Case - Data File

<p>Data file&nbsp;in CSV format of &quot;Perceptions Journal&quot; articles&nbsp; (2010-2012)</p>

opencc-by-4.0Jun 2018View details →
zenodo36/100

alawinia/provClustering: Discovering Similar Workflows via Provenance Clustering

<p>Several workflow management systems and scripting languages have adopted provenance tracking, yet many researchers choose to manually capture or instrument their processing scripts to write provenance information to files. The Next Generation Sequencing (NGS) project we are associated with is tracking provenance in such manner. The NGS project is a collaboration between multiple groups at different sites, where each group is collecting and processing samples using an agreed-upon workflow. The workflow contains many stages with varying degrees of complexity. Over time workflow stages are modified, but data samples are only comparable when processed with identical versions of the workflow. However, for various reasons (including the distributed nature of the collaboration) it is not always clear which samples have been processed with which version of the workflow. In this paper, we introduce new techniques for clustering provenance datasets and attempt to discover the ones that are likely to be generated by same workflow. Based on the clustering result, users can identify similar provenance and would be able to categorize them into different clusters for debugging and zoom-in/zoom-out viewing.</p>

openother-openJul 2018View details →
zenodo36/100

Rajapuri, Satara dst., Maharashtra. Cluster of hero-stones

<p>Cluster of hero-stones nearby Shri Kartik Swami shrine.</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Rajapuri, Satara dst., Maharashtra. Cluster of hero-stones (cave 3)

<p>Cluster of hero-stones, cave 3, Shri Kartik Swami complex</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Rajapuri, Satara dst., Maharashtra. Cluster of sati-stones (cave 1)

<p>Cluster of sati-stones, cave 1, Shri Kartik Swami complex</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Simulation dataset for "Computational pan-genome mapping and pairwise SNP-distance improve detection of Mycobacterium tuberculosis transmission clusters"

<p>Simulated Illumina reads for SNP distance method evaluation and comparison used in the article &quot;Computational pan-genome mapping and pairwise SNP-distance improve detection of Mycobacterium tuberculosis transmission clusters&quot;.</p> <p>Details for simulation can be found at https://gitlab.com/rki_bioinformatics/panpasco/tree/master/simulation_dataset.</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Properties of negative initial leaders and lightning flash size in a cluster of supercells

<p>It is the figure data of the paper titled as &quot;Properties of negative initial leaders and lightning flash size in a cluster of supercells&quot;. The abstract of this paper is as follows:</p> <p>Properties of negative initial leaders (NILs) and flash size in a cluster of supercells with generally inverted charge structure in Oklahoma on 10 &minus; 11 May 2010 are examined, primarily using Lightning Mapping Array data. A method to identify NILs from LMA source is proposed, and helps to reveal the multiple NILs properties and their distributions. The NILs in the supercell cluster have smaller speed (median 3D displacement speed: 0.65 &times; 10<sup>5</sup> m s<sup>&minus;1</sup>), relative to the previous reports in &ldquo;normal&rdquo; thunderstorms. Furthermore, median NIL speeds initially decrease with increasing height, but begin increasing above 12 km. The NILs tend to decelerate during the early stage. The parameters characterizing flash duration and spatial size are also investigated. It is found they all follow lognormal distributions and the spatial flash size is relatively small on average (median horizontal distance: 5.54 km). Most flashes (83.18%) extend primarily in the horizontal direction. Flash area shows an inverse relationship with flash density at their fast changes during storm evolution. Although large flash initiation density (FID) generally occurs in regions with small flash size, the smallest flash size is nearly not collocated with large FID value. In the regions with large FID, average flash duration roughly increases with increasing FID, while average flash area changes little. We proposed that the pattern of charge pockets and variation of charge density dominated by the strong kinematics are responsible for some new findings about the properties of NIL and flash size in the supercell cluster.</p>

opencc-by-4.0Sep 2018View details →
zenodo36/100

On the Observability of Individual Population III Stars and Their Stellar-mass Black Hole Accretion Disks through Cluster Caustic Transits

<p>MESA inlists associated with&nbsp;<a href="https://ui.adsabs.harvard.edu/#abs/2018ApJS..234...41W/abstract">On the Observability of Individual Population III Stars and Their Stellar-mass Black Hole Accretion Disks through Cluster Caustic Transits</a></p>

opencc-by-4.0Mar 2019View details →
zenodo36/100

Formation of Black Hole X-Ray Binaries with Non-degenerate Donors in Globular Clusters

<p>MESA inlists associated with&nbsp;<a href="https://ui.adsabs.harvard.edu/?#abs/2017ApJ...843L..30I">Formation of Black Hole X-Ray Binaries with Non-degenerate Donors in Globular Clusters</a></p>

opencc-by-4.0Mar 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record