Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18
datasets available to search
ShareScore release 0.9.0
Dataset results
18 results for “clustering algorithms”
Application of two-step clustering algorithm to QuaLiKiz-v2.6.2 turbulent transport simulation data
<p>QuaLiKiz simulation data in support of the two-step clustering algorithm, developed by Bart J. J. Kremers.</p> <p>The NETCDF file, generated via NETCDF4, contains the raw QuaLiKiz output for the 3-dimensional (2-input, 1-output) toy case used to develop the algorithm. Within the NETCDF file, the coordinates represent the code inputs and various vector indices and the data variables represent the code outputs.</p> <p>There are also 4 HDF5 files, containing the results from the two-step clustering reduction algorithm, where the data is saved under 2 keys: "/input" and "/flattened". The file names indicate the reduction algorithm settings used to produce the results within.</p> <p>The algorithm is available open-source at <a href="https://gitlab.com/BartKremers/two-step-clustering">https://gitlab.com/BartKremers/two-step-clustering</a>.</p>
Sensitivity and cluster by spectral clustering algorithm
<p>File contains the node pressure sensitivity matrix to 4LPS leaks and the cluster matrix for each node determined using the spectral clustering algorithm.</p>
Benne: A Modular Data Stream Clustering Algorithm with Flexible Design Choices
<p>All of the source dataset with preprocessed format [id features class] that have been used for evaluation in the paper.</p>
Figure 3: AntCo2 algorithm for graph clustering: on the left the output of the computation on a communication network; on the right the output on a regular grid
<p>Social and human developments are typical complex systems. Urban development<br> and dynamics are the perfect illustration of systems where spatial<br> emergence, self-organization and structural interaction between the system<br> and its components occur [3, 4, 5, 6]. In figure 4, we concentrate on the emergence<br> of organizational systems from geographical systems.</p>
Figure 2: Optimization in natural ants collective behavior: foraging and clustering (from [8])-Self-organization and social insects algorithms
<p>On figure 2, two examples of self-organization in natural ants are presented.<br> On the left side, the well-known Deneubourg experiment consists to highlight<br> with a very simple device the ant foraging problem. The ant objectives is<br> to find the optimal way from nest to food source, using pheromone trail deposition.<br> On the right side, cemetery clustering formation are shown at 4<br> successive times: ants form piles of corpses to clean their nests. Each of them<br> has elementary actions, unknowing the whole situation, but dealing only with<br> local information. There is no supervisor to lead the piles formation which<br> emerges from ant interactions.</p>
Data from: Algorithm for post-clustering curation of DNA amplicon data yields reliable biodiversity estimates
DNA metabarcoding is promising for cost-effective biodiversity monitoring, but reliable diversity estimates are difficult to achieve and validate. Here we present and validate a method, called LULU, for removing erroneous molecular operational taxonomic units (OTUs) from community data derived by high-throughput sequencing of amplified marker genes. LULU identifies errors by combining sequence similarity and co-occurrence patterns. To validate the LULU method, we use a unique data set of high quality survey data of vascular plants paired with plant ITS2 metabarcoding data of DNA extracted from soil from 130 sites in Denmark spanning major environmental gradients. OTU tables are produced with several different OTU definition algorithms and subsequently curated with LULU, and validated against field survey data. LULU curation consistently improves α-diversity estimates and other biodiversity metrics, and does not require a sequence reference database; thus, it represents a promising method for reliable biodiversity estimation.
A supervised Graph-based deep learning algorithm to detect and quantify clustered particles
<p>In this data repository, we provide the necessary data for replicating results, including both simulated and biological datasets. Additionally, the repository includes trained models to infer from these datasets.</p>
Benchmark cancer datasets for Clustering algorithms for Omics-based Patient Stratification (COPS)
<p>This repository contains seven multi-omic cancer datasets including several cancer types (breast, kidney, lung, ovary, prostate, and thyroid cancers as well as low grade gliomas) that were used for benchmarking several multi-view clustering algorithms implemented by COPS (https://github.com/UEFBiomedicalInformaticsLab/COPS). The datasets were originally compiled from The Cancer Genoma Atlas (TCGA) and downloaded using the <em>curatedTCGAData</em> R-package. The datasets include copy-number variations, methylomics as well as mRNA and miRNA transcriptomics. The methylomics data was mapped to genes by averaging methylation level of probes associated with the promoter regions of genes. Similarly the miRNA transcriptomics data was mapped to genes by using known and predicted miRNA -> gene interactions. Updated survival data was acquired from the Liu et al. 2018 paper. </p> <p>This repository also includes two sets of cancer associated pathway networks used by pathway-based multi-omic methods benchmarked in our study. NCI-PID pathways were downloaded using the <em>ndexr</em> R-package on December 22 2021. While KEGG pathways were downloaded using the <em>pathview</em> R-package on May 3 2022. </p> <p>More details on the processing can be found on the related publication.</p>
Initial application of the noise-sorted scanning clustering algorithm to the analysis of composition-dependent organic aerosol thermal desorption measurements
Open the record for dataset details and reuse information.
Data from: Algorithm for post-clustering curation of DNA amplicon data yields reliable biodiversity estimates
Open the record for dataset details and reuse information.
Structural files of algorithmically generated molybdenum oxide clusters
<p>Structural files of algorithmically generated molybdenum oxide clusters. The new clusters were generated to fit experimental pair distribution functions of molybdenum oxide surface layers supported on Al<sub>2</sub>O<sub>3 </sub>and zeolites. Results are published in article XXX</p>
Multiple Sclerosis lesions detection by a hybrid Watershed-Clustering algorithm
<p>Computer Aided Diagnosis (CAD) systems have been developing in the last years with the aim of helping the diagnosis and monitoring of several diseases. We present a novel CAD system based on a hybrid Watershed-Clustering algorithm for the detection of lesions in Multiple Sclerosis. Magnetic Resonance Imaging scans (FLAIR sequences without gadolinium) of 20 patients affected by Multiple Sclerosis with hyperintense lesions were studied. The CAD system consisted of the following automated processing steps: images recording, automated segmentation based on the Watershed algorithm, detection of lesions, extraction of both dynamic and morphological features, and classification of lesions by Cluster Analysis. The investigation was performed on 316 suspect regions including 255 lesion and 61 non-lesion cases. The Receiver Operating Characteristic analysis revealed a highly significant difference between lesions and non-lesions; the diagnostic accuracy was 87% (95% CI: 0.83–0.90), with an appropriate cut-off of 192.8; the sensitivity was 77% and the specificity was 87%. In conclusion, we developed a CAD system by using a modified algorithm for automated image segmentation which may discriminate MS lesions from non-lesions. The proposed method generates a detection out-put that may be support the clinical evaluation.</p>
A real dataset for evaluating hybrid clustering algorithm
<p>This is a real dataset for testing our hybrid clustering algorithm, and details can be found in our manuscript.</p>
Datasets S1 ~ S6 for evaluating hybrid clustering algorithm
<p>These datasets are used to evaluate our hybrid clustering algorithm. For details of our algorithm, please refer to https://github.com/junhaiqi/Hybrid_clustering.git.</p>
Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis
The inference of population genetic structures is essential in many research areas in population genetics, conservation biology and evolutionary biology. Recently, unsupervised Bayesian clustering algorithms have been developed to detect a hidden population structure from genotypic data, assuming among others that individuals taken from the population are unrelated. Because of this hypothesis, markers in a sample taken from a subpopulation can be considered to be in Hardy-Weinberg and linkage equilibrium. However, close relatives might be sampled from the same subpopulation, and consequently, might cause Hardy-Weinberg and linkage disequilibrium and thus bias a population genetic structure analysis. In this study, we used simulated and real data to investigate the impact of close relatives in a sample on Bayesian population structure analysis. We also showed that, when close relatives were identified by a pedigree reconstruction approach and removed, the accuracy of a population genetic structure analysis can be greatly improved. The results indicate that unsupervised Bayesian clustering algorithms cannot be used blindly to detect genetic structure in a sample with closely related individuals. Rather, when closely related individuals are suspected to be frequent in a sample, these individuals should be first identified and removed before conducting a population structure analysis.
Data from: The effect of close relatives on unsupervised Bayesian clustering algorithms in population genetic structure analysis
Open the record for dataset details and reuse information.
DM-PhyClus: A Bayesian phylogenetic algorithm for infectious disease transmission cluster inference
<p>The files include the simulated datasets and the chain results for the simulation study described in, </p> <p>DM-PhyClus: A Bayesian phylogenetic algorithm for infectious disease transmission cluster inference</p> <p>Those are all compressed R data files (gzip format).</p>
Geostationary Lightning Mapper (GLM) Cluster Integrity, Exception Resolution, and Reclustering Algorithm (CIERRA)
The Geostationary Lightning Mapper (GLM) Cluster Integrity, Exception Resolution, and Reclustering Algorithm (CIERRA) dataset consists of a hierarchy of earth-located lightning radiant energy measures including events, groups, series, flashes, and areas. The GLM CIERRA data addresses the artificial flash termination by the GLM ground system by recombining split flashes and filtering out more non-lightning noise. This provides researchers with a powerful tool to better investigate convective storm and lightning activity with more accurate observations as well as better incorporate spatial extent observations that can be used for aviation meteorology, lightning safety, and other studies. These data are available from January 12, 2017, through March 31, 2023, in netCDF-4 format.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.