Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,326
datasets available to search
ShareScore release 0.9.0
Dataset results
2,326 results for “clusters”
Cataloging Distant Galactic Open Clusters: Identification of 739 New Star Clusters Beyond 5 kpc Utilizing GAIA DR3 Data
<p><span>The figures of the 739 open clusters that report in our paper (Cataloging Distant Galactic Open Clusters: Identification of 739 New Star Clusters Beyond 5 kpc Utilizing GAIA DR3 Data)</span></p> <p><span> complete King's model<span> </span>profile fitting, sky charts and 5-panels ( spatial distribution, proper-motion distribution,</span></p> <p><span><span>parallax statistics, parallax distribution, and CMD)</span> .</span></p>
Code and data for spatial and temporal magnitude clustering analysis
<p>Code used for performing spatial and temporal seismic magnitude clustering analysis. Includes documentation (README.txt) with steps on how to implement the code. The public datasets used for this study can be accessed at the following locations: </p> <ul> <li><strong>Southern California Catalog: </strong> <ul> <li>SCEDC (2013): Southern California Earthquake Center.<br> Caltech.Dataset. doi:<a href="https://dx.doi.org/10.7909/C3WD3xH1">10.7909/C3WD3xH1</a></li> </ul> </li> <li><strong>Northern California Catalog:</strong> <ul> <li>NCEDC (2014), Northern California Earthquake Data Center. UC Berkeley Seismological Laboratory. Dataset. doi:10.7932/NCEDC.</li> </ul> </li> <li><strong>Mixed-mode Laboratory Catalog:</strong> <ul> <li>Lin, Qing, et al. "Opening and mixed mode fracture processes in a quasi-brittle material via digital imaging." <em>Engineering Fracture Mechanics</em> 131 (2014): 176-193.</li> </ul> </li> <li><strong>ETAS Code:</strong> <ul> <li>Leila Mizrahi, Shyam Nandan, Stefan Wiemer 2021;<br> Embracing Data Incompleteness for Better Earthquake Forecasting. (Section 3.1)<br> <em>Journal of Geophysical Research: Solid Earth</em>; doi: <a href="https://doi.org/10.1029/2021JB022379">https://doi.org/10.1029/2021JB022379</a></li> </ul> </li> </ul>
Clust&See3.0 : clustering, module exploration and annotation.
<p><span>Clust&See</span><span>3</span><span>.0 is the novel version of a </span><span>Cytoscape</span><span> app </span><span>that has been developed to </span><span>identify, visualize and manipulate network clusters and modules, newly enriched with functionalities allowing </span><span>custom </span><span>annotation</span><span>s</span><span> of nodes and computation of their statistical enrichment</span><span>s</span><span>. As the wealth of multi-omics data is growing,</span><span> such functionalities are highly valuable for a better understanding of biological module composition</span><span>.</span></p> <p> </p> <ul> <li><span>go_annot.txt : The protein annotation file extracted from Gene Ontology Biological Process database.</span></li> <li><span>HuRI_CC.txt : The network file containing the largest connect component of <span>the </span>human reference interactome network [1]</span></li> </ul> <p><span>[1] <span>Luck K, Kim D-K, Lambourne L, Spirohn K, Begg BE, Bian W, et al. A reference map of the human binary protein interactome. Nature. 2020;580:402–8. </span></span></p>
Deciphering the explanatory potential of blood pressure variables on post-operative length of stay through hierarchical clustering: A retrospective monocentric study
<p><em>Objective:</em> Mean arterial pressure is widely used as the variable to monitor during anesthesia. But there are many other variables proposed to define intraoperative arterial hypotension. The goal of the present study was to search arterial pressure variables linked with prolonged postoperative length of stay (pLOS).</p> <p><em>Design: </em>Retrospective cohort study of adult patients having received general for a scheduled non cardiac surgical procedure between 15<sup>th</sup> July 2017 and 31st December 2019.</p> <p><em>Methods:</em> pLOS was defined as a stay longer than the median (main outcome), adjusted for surgery type and duration. 330 arterial pressure variables were analyzed and organized through a clustering approach. An unsupervised hierarchical aggregation method for optimal cluster determination, employing Kendall's tau coefficients and a penalized Bayes information criterion was used. Variables were ranked using the absolute standardized mean distance (aSMD) to measure their effect on pLOS. Finally, after multivariate independence analysis, the number of variables was reduced to three.</p> <p><em>Results:</em> Our study examined 9,516 patients. When LOS is defined as strictly greater than the median, 34% of patients experienced pLOS. Key arterial pressure variables linked with this definition of pLOS included the difference between the highest and lowest pulse pressure values computed throughout the surgery (aSMD[95%CI] =0.39[0.31-0.40], p<0.001), the accumulated time pulse pressure above 61mmHg (aSMD = 0.21[0.17-0.25], p<0.001), and the lowest MAP during surgery (aSMD= 0.20[0.16-0.24], p<0.001).</p> <p><em>Conclusions: </em>By applying a clustering approach, three arterial pressure variables were associated with pLOS. This scalable method can be applied to various dichotomized outcomes.</p>
Clustering Tasks and Decision Trees with Elegiac Poets
<p>The dataset contains files generated during a Natural Language Processing (NLP) and automatic text analysis task. Attached is a <strong>Jupyter notebook</strong> with the complete code, along with several <strong>Excel files (.xlsx)</strong> containing organized information. Additionally, there are three folders that include files generated during the Silhouette calculation, K-means clustering, and feature extraction using decision trees.</p> <p>The three folders are:<br>1. <strong>Silhouette Calculation:</strong> Contains PNG images of Silhouette plots for various analysis configurations.<br>2.<strong> K-means Clustering: </strong>Contains pickle (.pkl) files with features and labels for each combination of excluded author, n-gram type, n-gram range, and matrix type.<br>3. <strong>Feature Extraction: </strong>Contains CSV files with lists of documents by cluster and the most important features along with information gain and information gain ratio metrics.</p> <p>Other file formats included in the dataset are:<br>- CSV files containing Silhouette scores, optimal clustering results, cluster assignments, and optimal cluster assignments.<br>- PNG images of scatter plots colored by author and by cluster.<br>- Pickle files containing the top features extracted during the analysis.</p>
Figure 3. Cluster Dendrogram indicating 11 in Composition and structure of plant communities in the Moist Temperate Forest Ecosystem of the Hindukush Mountains, Pakistan
Figure 3. Cluster Dendrogram indicating 11 plant association types in the Lalkoo Valley.
Benchmark cancer datasets for Clustering algorithms for Omics-based Patient Stratification (COPS)
<p>This repository contains seven multi-omic cancer datasets including several cancer types (breast, kidney, lung, ovary, prostate, and thyroid cancers as well as low grade gliomas) that were used for benchmarking several multi-view clustering algorithms implemented by COPS (https://github.com/UEFBiomedicalInformaticsLab/COPS). The datasets were originally compiled from The Cancer Genoma Atlas (TCGA) and downloaded using the <em>curatedTCGAData</em> R-package. The datasets include copy-number variations, methylomics as well as mRNA and miRNA transcriptomics. The methylomics data was mapped to genes by averaging methylation level of probes associated with the promoter regions of genes. Similarly the miRNA transcriptomics data was mapped to genes by using known and predicted miRNA -> gene interactions. Updated survival data was acquired from the Liu et al. 2018 paper. </p> <p>This repository also includes two sets of cancer associated pathway networks used by pathway-based multi-omic methods benchmarked in our study. NCI-PID pathways were downloaded using the <em>ndexr</em> R-package on December 22 2021. While KEGG pathways were downloaded using the <em>pathview</em> R-package on May 3 2022. </p> <p>More details on the processing can be found on the related publication.</p>
Galaxy populations in the Hydra I cluster from the VEGAS survey III. The realm of low-surface brightness features and intra-cluster light
<p><span>This appendix provides the azimuthally-averaged surface bright</span><span>ness and colour profiles of the sample galaxies listed in Table 1 of the paper "Galaxy populations in the Hydra I cluster from the VEGAS survey III. The realm of low-surface brightness features and intra-cluster light", by <span>Marilena Spavone</span><span>,</span><span> Enrichetta Iodice</span><span>,</span><span> Felipe S. Lohmann</span><span>, Magda Arnaboldi</span><span>, Michael Hilker</span><span>, Antonio La </span><span>Marca</span><span>, Rosa Calvi</span><span>, Michele Cantiello</span><span>, Enrico M. Corsini</span><span>, Giuseppe D’Ago</span><span>, Duncan A. Forbes</span><span>, Marco </span><span>Mirabile</span><span>, and Marina Rejkuba.</span><br></span></p>
Data file needed for star cluster setup in Phantom smoothed particle hydrodynamics and magnetohydrodynamics code
<p>** this file is automatically downloaded by Phantom when running the starcluster setup **</p> <p>This is a small ascii file containing positions and velocities of stars utilised in the "starcluster" configuration in the Phantom smoothed particle hydrodynamics and magnetohydrodynamics code (<a href="http://adsabs.harvard.edu/abs/2018PASA...35...31P">Price et al. 2018</a>). It is used to set up a collection of N-body particles.</p> <p>The star cluster setup (and the data file) were written by Yann Bernard as part of his PhD thesis at<strong> </strong>Université Grenoble Alpes. The datafile is published here so it can be used in the automated code testing via github actions. </p> <p>The columns are mass, position (x,y,z) and velocity (vx,vy,vz) for all of the stars in the simulation</p>
Original NGS dataset from publication "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center".
<p>Original hepatitis C virus hypervariable region 1 NGS sequences in fastq format from patients analyzed in the study "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center". </p> <p> </p>
Influence of temperature on the molecular composition of ions and charged clusters during pure biogenic nucleation
<p>Data of Figures in manuscript:</p> <p>Influence of temperature on the molecular composition of ions and charged clusters during pure biogenic nucleation.</p> <p>Frege et al., Atmos. Chem. Phys. 18, 65–79, 2018</p> <p>https://doi.org/10.5194/acp-18-65-2018</p> <p> </p>
International Relations Articles Clustering Case - Data File
<p>Data file in CSV format of "Perceptions Journal" articles (2010-2012)</p>
alawinia/provClustering: Discovering Similar Workflows via Provenance Clustering
<p>Several workflow management systems and scripting languages have adopted provenance tracking, yet many researchers choose to manually capture or instrument their processing scripts to write provenance information to files. The Next Generation Sequencing (NGS) project we are associated with is tracking provenance in such manner. The NGS project is a collaboration between multiple groups at different sites, where each group is collecting and processing samples using an agreed-upon workflow. The workflow contains many stages with varying degrees of complexity. Over time workflow stages are modified, but data samples are only comparable when processed with identical versions of the workflow. However, for various reasons (including the distributed nature of the collaboration) it is not always clear which samples have been processed with which version of the workflow. In this paper, we introduce new techniques for clustering provenance datasets and attempt to discover the ones that are likely to be generated by same workflow. Based on the clustering result, users can identify similar provenance and would be able to categorize them into different clusters for debugging and zoom-in/zoom-out viewing.</p>
Rajapuri, Satara dst., Maharashtra. Cluster of hero-stones
<p>Cluster of hero-stones nearby Shri Kartik Swami shrine.</p>
Rajapuri, Satara dst., Maharashtra. Cluster of hero-stones (cave 3)
<p>Cluster of hero-stones, cave 3, Shri Kartik Swami complex</p>
Rajapuri, Satara dst., Maharashtra. Cluster of sati-stones (cave 1)
<p>Cluster of sati-stones, cave 1, Shri Kartik Swami complex</p>
Simulation dataset for "Computational pan-genome mapping and pairwise SNP-distance improve detection of Mycobacterium tuberculosis transmission clusters"
<p>Simulated Illumina reads for SNP distance method evaluation and comparison used in the article "Computational pan-genome mapping and pairwise SNP-distance improve detection of Mycobacterium tuberculosis transmission clusters".</p> <p>Details for simulation can be found at https://gitlab.com/rki_bioinformatics/panpasco/tree/master/simulation_dataset.</p>
Properties of negative initial leaders and lightning flash size in a cluster of supercells
<p>It is the figure data of the paper titled as "Properties of negative initial leaders and lightning flash size in a cluster of supercells". The abstract of this paper is as follows:</p> <p>Properties of negative initial leaders (NILs) and flash size in a cluster of supercells with generally inverted charge structure in Oklahoma on 10 − 11 May 2010 are examined, primarily using Lightning Mapping Array data. A method to identify NILs from LMA source is proposed, and helps to reveal the multiple NILs properties and their distributions. The NILs in the supercell cluster have smaller speed (median 3D displacement speed: 0.65 × 10<sup>5</sup> m s<sup>−1</sup>), relative to the previous reports in “normal” thunderstorms. Furthermore, median NIL speeds initially decrease with increasing height, but begin increasing above 12 km. The NILs tend to decelerate during the early stage. The parameters characterizing flash duration and spatial size are also investigated. It is found they all follow lognormal distributions and the spatial flash size is relatively small on average (median horizontal distance: 5.54 km). Most flashes (83.18%) extend primarily in the horizontal direction. Flash area shows an inverse relationship with flash density at their fast changes during storm evolution. Although large flash initiation density (FID) generally occurs in regions with small flash size, the smallest flash size is nearly not collocated with large FID value. In the regions with large FID, average flash duration roughly increases with increasing FID, while average flash area changes little. We proposed that the pattern of charge pockets and variation of charge density dominated by the strong kinematics are responsible for some new findings about the properties of NIL and flash size in the supercell cluster.</p>
On the Observability of Individual Population III Stars and Their Stellar-mass Black Hole Accretion Disks through Cluster Caustic Transits
<p>MESA inlists associated with <a href="https://ui.adsabs.harvard.edu/#abs/2018ApJS..234...41W/abstract">On the Observability of Individual Population III Stars and Their Stellar-mass Black Hole Accretion Disks through Cluster Caustic Transits</a></p>
Formation of Black Hole X-Ray Binaries with Non-degenerate Donors in Globular Clusters
<p>MESA inlists associated with <a href="https://ui.adsabs.harvard.edu/?#abs/2017ApJ...843L..30I">Formation of Black Hole X-Ray Binaries with Non-degenerate Donors in Globular Clusters</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.