Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
155
datasets available to search
ShareScore release 0.9.0
Dataset results
155 results for “Cluster analysis”
Fig. 2 in Cluster Analysis of Non-conserved Proteins of Trypanosoma cruzi Reference Strains Displays Parity between these Groupings (Peptidemes) and the Consensually Accepted Parasite Lineages
Fig. 2. Diagrammatic representation of the twenty-two protein bands not shared by all Trypanosoma cruzi reference strains (nonconserved proteins), as visualized in SDS-PAGE. These bands were coded and analyzed by numerical taxonomy procedures. At the top is indicated the number of the major groups they belong, as identified by different approaches. The bands that were exclusive of one or more strains were highlighted with rectangles. M: molecular mass markers. (kDa) are indicated on the left.
Fig. 2. Hierarchical cluster analysis with 2 in Proliferation of the invasive termite Coptotermes gestroi (Isoptera: Rhinotermitidae) on Grand Cayman and overall termite diversity on the Cayman Islands
Fig. 2. Hierarchical cluster analysis with 2 (a), 3 (b), 4 (c), and 5 (d) clusters for Coptotermes gestroi over Grand Cayman Island.
BRAIN Journal-Performance Analysis of Unsupervised Clustering Methods for Brain Tumor Segmentation-Figure 4.Performance based on no. of tumor pixel & execution time
<p>In this paper we segmented the brain tumors in axial view of MR images with the help of<br> unsupervised clustering method i.e. K-means clustering. The unsupervised clustering methods gave<br> the better results than traditional method.<br> The performance analysis and comparison is done f on the basis of no. of tumor pixels in<br> segmented brain tumor and the execution time for the same. Regarding the no. of tumor pixels, Kmeans<br> clustering gave a better result than the other methods. The clustering algorithms were tested<br> with a data base of 20 MRI brain images. K-means clustering achieved almost 90%result</p>
BRAIN Journal-Performance Analysis of Unsupervised Clustering Methods for Brain Tumor Segmentation-Figure 3:(a) Input MR Image (b) Enhanced Image (c) Segmented Tumor (d) Located brain tumor
<p>Figure 3 shows three different original brain MR images, contrast enhancement of the<br> images, segmented images using K-means algorithm and finally located tumor. Fig 1.4 shows the<br> performance of the unsupervised clustering methods with the no. of tumor pixels and execution<br> time to locate the brain tumor.</p>
BRAIN Journal-Performance Analysis of Unsupervised Clustering Methods for Brain Tumor Segmentation-Figure 2. Stages of software implementation
<p>The algorithm has two stages, first is pre-processing of given MRI image and after that<br> segmentation and then perform morphological operations.</p>
BRAIN Journal-Performance Analysis of Unsupervised Clustering Methods for Brain Tumor Segmentation-Figure 1. Diagnosis Rate in different Countrie
<p>In MRI images, the amount of data is too much for manual segmentation. The procedure is<br> tedious, time, labor consuming, subjective and requires expertise. This gave way to methods that are<br> computer-aided with user interaction at varying levels. These methods are automatic and objective<br> and the results are highly reproducible. We designed software tool for locating brain tumor, based<br> on unsupervised clustering methods and analyzed its performance</p>
Figure 9. Performance analysis of FCM-PSO, GPC-PSO and GFCM-PSO-An Optimized Clustering Approach for Automated Detection of White Matter Lesions in MRI Brain Images
<p>All scans obtained from different image clustering models are manually ranked based on<br> values in table 1. Table 2 represents WML detection rates of optimized images. FCM, GPC and<br> GFCM clustering methods and hybrid optimized methods (FCM-PSO, GPC-PSO and GFCM-PSO)<br> are applied on a dataset of 208 images and ranking is done in terms of under detected, over<br> detected, properly detected as shown in figure 8 and figure 9.</p>
Figure 8. Performance analysis of FCM, GPC and GFCM Figure 9.-An Optimized Clustering Approach for Automated Detection of White Matter Lesions in MRI Brain Images
<p>All scans obtained from different image clustering models are manually ranked based on<br> values in table 1. Table 2 represents WML detection rates of optimized images. FCM, GPC and<br> GFCM clustering methods and hybrid optimized methods (FCM-PSO, GPC-PSO and GFCM-PSO)<br> are applied on a dataset of 208 images and ranking is done in terms of under detected, over<br> detected, properly detected as shown in figure 8 and figure 9. The number of images detected<br> properly in GFCM is comparatively high than FCM and GPC. The optimized result of GFMC<br> provides accurate detection of WMLs and it properly detects 195 images.</p>
Fig. 4 – Cluster analysis results for the distance from which ants mobilize from the nest. 4 in Mobilization Strategies in Ants (Hymenoptera: Formicidae)
Fig. 4 – Cluster analysis results for the distance from which ants mobilize from the nest. 4th cluster: F cin – Formica cinerea; 3rd cluster: F. fus – F. fusca; M. rub – Myrmica rubra; M. mut – Messor muticus; M. rug – Myrmica ruginodis; L. bru – Lasius brunneus; D. qua – Dolichoderus quadripunctatus; L. nig – Lasius niger; L. pla – L. platythorax; L. neg U – L. neglectus; T. cae – Tetramorium caespitum; T. err – Tapinoma erraticum; Temn – Temnothorax sp.; P. pal – Plagiolepis pallescens; L. ace – Leptothorax acervorum; S. fug – Solenopsis fugax; P. tau – Plagiolepis tauricus; T. arm – Tetramorium armatum; M. sal – Myrmica salina; M. spe – M. specioides; F. cun – Formica cunicularia; L. ema – Lasius emarginatus; F. ruf – Formica rufibarbis; F. cla – F. clara; C. aet – Camponotus aethiops; 2nd cluster: F. tru – Formica truncorum; L. ful – Lasius fuliginosus; C. vag – Camponotus vagus; F. pra – Formica pratensis; 1st cluster: F. rufa – Formica rufa; F. pol – F. polyctena; C. sub U – Crematogaster subdentata.
Extended data tables to Haering and Habermann, F1000Res, RNfuzzyApp: an R shiny RNA-seq data analysis app for visualisation, differential expression analysis, time-series clustering and enrichment analysis
<p><b>Background</b> </p> <p>RNA-seq is a widely adopted affordable method for large scale gene expression profiling. However, user-friendly and versatile tools for wet-lab biologists to analyse RNA-seq data beyond standard analyses such as differential expression, are rare. Especially, the analysis of time-series data is difficult for wet-lab biologists lacking advanced computational training. Furthermore, most meta-analysis tools are tailored for model organisms and not easily adaptable to other species.</p> <p><b>Results</b></p> <p>With RNfuzzyApp, we provide a user-friendly, web-based R-shiny app for differential expression analysis, as well as time-series analysis of RNA-seq data. RNfuzzyApp offers several methods for normalization and differential expression analysis of RNA-seq data, providing easy-to-use toolboxes, interactive plots and downloadable results. For time-series analysis, RNfuzzyApp presents the first web-based, automated pipeline for soft clustering with the Mfuzz R package, including methods to aid in cluster number selection, Mfuzz loop computations, cluster overlap analysis, as well as cluster enrichments.</p> <p><b>Conclusion</b></p> <p>RNfuzzyApp is an intuitive, easy to use and interactive R shiny app for RNA-seq differential expression and time-series analysis, offering a rich selection of interactive plots, providing a quick overview of raw data and generating rapid analysis results. Furthermore, its orthology assignment, enrichment analysis, as well as ID conversion functions are accessible to non-model organisms.</p>
Text-fig. 12. Scanning electron microscope (SEM) images of pollen of Sergipea sp. from a group of probable fragmentary pollen sacs; Torres Vedras locality, Portugal. a) Cluster of probable fragmentary pollen sacs that yielded the pollen in this Text-figure; b, c) Pollen grains showing the robust longitudinal ribs separated by prominent areas of granular exine; note the groove along the margins of the longitudinal ribs (arrowheads); d) Pollen grain showing the granular exine flanked by two robust ribs; note the groove along the margins of the longitudinal ribs (arrowheads). Specimen, TV44-S148012 (a–d). Scale bars 150 Μm (a), 12 Μm (c), 6 Μm (b, d). in The Early Cretaceous Mesofossil Flora Of Torres Vedras (Ne Of Forte Da Forca), Portugal: A Palaeofloristic Analysis Of An Early Angiosperm Community
Text-fig. 12. Scanning electron microscope (SEM) images of pollen of Sergipea sp. from a group of probable fragmentary pollen sacs; Torres Vedras locality, Portugal. a) Cluster of probable fragmentary pollen sacs that yielded the pollen in this Text-figure; b, c) Pollen grains showing the robust longitudinal ribs separated by prominent areas of granular exine; note the groove along the margins of the longitudinal ribs (arrowheads); d) Pollen grain showing the granular exine flanked by two robust ribs; note the groove along the margins of the longitudinal ribs (arrowheads). Specimen, TV44-S148012 (a–d). Scale bars 150 Μm (a), 12 Μm (c), 6 Μm (b, d).
Text-fig. 18. Scanning electron microscope (SEM) images of a fruit of Canrightia sp. with associated pollen; Torres Vedras locality, Portugal. a) Fruit in lateral view showing prominent cavities in the fruit wall formed by the scattered oil bodies and the broad hypanthium fused to the base of the fruit (arrowhead); b) Fruit surface showing epidermal cells and the scattered oil cells embedded in the fruit wall (arrowheads); c) Cluster of monocolpate pollen grains in the probable stigmatic region of the fruit; d) Pollen grains showing the long colpus and semitectate-reticulate pollen wall; e) Pollen wall showing the reticulum with large and small lumina, and scattered, compressed columellae supporting the smooth muri. Specimen, TV142-S170213. Scale bars 300 Μm (a), 100 Μm (b), 30 Μm (c), 6 Μm (d), 1 Μm (e). in The Early Cretaceous Mesofossil Flora Of Torres Vedras (Ne Of Forte Da Forca), Portugal: A Palaeofloristic Analysis Of An Early Angiosperm Community
Text-fig. 18. Scanning electron microscope (SEM) images of a fruit of Canrightia sp. with associated pollen; Torres Vedras locality, Portugal. a) Fruit in lateral view showing prominent cavities in the fruit wall formed by the scattered oil bodies and the broad hypanthium fused to the base of the fruit (arrowhead); b) Fruit surface showing epidermal cells and the scattered oil cells embedded in the fruit wall (arrowheads); c) Cluster of monocolpate pollen grains in the probable stigmatic region of the fruit; d) Pollen grains showing the long colpus and semitectate-reticulate pollen wall; e) Pollen wall showing the reticulum with large and small lumina, and scattered, compressed columellae supporting the smooth muri. Specimen, TV142-S170213. Scale bars 300 Μm (a), 100 Μm (b), 30 Μm (c), 6 Μm (d), 1 Μm (e).
A niching particle swarm optimization strategy combined with cluster analysis for the multimodal inversion of surface waves
<p>The data include two study cases used for multimodal surface wave inversion.</p> <p>For case 1, the data present a combination of active and passive surface wave methods.</p> <p>For case 3, we use Rayleigh waves to detect a low-velocity soft interlayer underneath the road.</p> <p>Detailed description can be found in the data description document.</p>
PhageHostLearn - training data and cluster analysis
<p>These data comprise the processed phage RBP and <em>Klebsiella </em>K-loci sequence data to train our PhageHostLearn system, along with ESM-2 embeddings of the RBPs and loci, as well as results from the cluster analyses of K-loci proteins and RBPs at 50% identity with CD-HIT.</p>
Upscaling soil organic carbon measurements at the continental scale using multivariate clustering analysis and machine learning
<p><strong>Data Description</strong>:</p> <p>To improve SOC estimation in the United States, we upscaled site-based SOC measurements to the continental scale using multivariate geographic clustering (MGC) approach coupled with machine learning models. First, we used the MGC approach to segment the United States at 30 arc second resolution based on principal component information from environmental covariates (gNATSGO soil properties, WorldClim bioclimatic variables, MODIS biological variables, and physiographic variables) to 20 SOC regions. We then trained separate random forest model ensembles for each of the SOC regions identified using environmental covariates and soil profile measurements from the International Soil Carbon Network (ISCN) and an Alaska soil profile data. We estimated United States SOC for 0-30 cm and 0-100 cm depths were 52.6 + 3.2 and 108.3 + 8.2 Pg C, respectively.</p> <p>Files in collection (32):</p> <p>Collection contains 22 soil properties geospatial rasters, 4 soil SOC geospatial rasters, 2 ISCN site SOC observations csv files, and 4 R scripts</p> <p>gNATSGO TIF files:</p> <p>├── available_water_storage_30arc_30cm_us.tif [30 cm depth soil available water storage]<br> ├── available_water_storage_30arc_100cm_us.tif [100 cm depth soil available water storage]<br> ├── caco3_30arc_30cm_us.tif [30 cm depth soil CaCO3 content]<br> ├── caco3_30arc_100cm_us.tif [100 cm depth soil CaCO3 content]<br> ├── cec_30arc_30cm_us.tif [30 cm depth soil cation exchange capacity]<br> ├── cec_30arc_100cm_us.tif [100 cm depth soil cation exchange capacity]<br> ├── clay_30arc_30cm_us.tif [30 cm depth soil clay content]<br> ├── clay_30arc_100cm_us.tif [100 cm depth soil clay content]<br> ├── depthWT_30arc_us.tif [depth to water table]<br> ├── kfactor_30arc_30cm_us.tif [30 cm depth soil erosion factor]<br> ├── kfactor_30arc_100cm_us.tif [100 cm depth soil erosion factor]<br> ├── ph_30arc_100cm_us.tif [100 cm depth soil pH]<br> ├── ph_30arc_100cm_us.tif [30 cm depth soil pH]<br> ├── pondingFre_30arc_us.tif [ponding frequency]<br> ├── sand_30arc_30cm_us.tif [30 cm depth soil sand content]<br> ├── sand_30arc_100cm_us.tif [100 cm depth soil sand content]<br> ├── silt_30arc_30cm_us.tif [30 cm depth soil silt content]<br> ├── silt_30arc_100cm_us.tif [100 cm depth soil silt content]<br> ├── water_content_30arc_30cm_us.tif [30 cm depth soil water content]<br> └── water_content_30arc_100cm_us.tif [100 cm depth soil water content]</p> <p>SOC TIF files:</p> <p>├──30cm SOC mean.tif [30 cm depth soil SOC]<br> ├──100cm SOC mean.tif [100 cm depth soil SOC]<br> ├──30cm SOC CV.tif [30 cm depth soil SOC coefficient of variation]<br> └──100cm SOC CV.tif [100 cm depth soil SOC coefficient of variation]</p> <p>site observations csv files:</p> <p>ISCN_rmNRCS_addNCSS_30cm.csv 30cm ISCN sites SOC replaced NRCS sites with NCSS centroid removed data</p> <p>ISCN_rmNRCS_addNCSS_100cm.csv 100cm ISCN sites SOC replaced NRCS sites with NCSS centroid removed data</p> <p><br> <strong>Data format</strong>:</p> <p>Geospatial files are provided in Geotiff format in Lat/Lon WGS84 EPSG: 4326 projection at 30 arc second resolution.</p> <p><strong>Geospatial projection</strong>: </p> <pre><code>GEOGCS["GCS_WGS_1984", DATUM["D_WGS_1984", SPHEROID["WGS_1984",6378137,298.257223563]], PRIMEM["Greenwich",0], UNIT["Degree",0.017453292519943295]] (base) [jbk@theseus ltar_regionalization]$ g.proj -w GEOGCS["wgs84", DATUM["WGS_1984", SPHEROID["WGS_1984",6378137,298.257223563]], PRIMEM["Greenwich",0], UNIT["degree",0.0174532925199433]] </code></pre> <p> </p>
Database and Syntax for Analysis of the Paper: "Effects of introducing the WHO Labour Care Guide on Caesarean section: a pragmatic, stepped-wedge, cluster randomized trial in India"
<p>The following files contains the information used to analyze the trial “Implementing the WHO Labour Care Guide to reduce the use of Caesarean section in four hospitals in India: a pragmatic, stepped wedge, cluster randomized pilot trial” in which it was hypothesized that the intervention would promote correct LCG use by these providers, changing their labour monitoring and management practices to align with WHO’s intrapartum recommendations. In turn, this could reduce overuse of Caesarean section, improve maternal and newborn outcomes, and enhance women’s care experiences. </p> <p>Two datafiles with extension “csv” are uploaded. The databased named “LCG Trial Women Database (transition period included).csv” is the database which contains the data of the recruited women in the trial. There is one row per women. The databased named “LCG Trial Neonates Database (transition period included).csv” is the database which contains the data of the neonates born from the recruited women. There is one row per neonate.</p> <p>The excel file “Data Dictionary LCG to Share.xlsx” is the data dictionary of the two databases. In the sheet named “Maternal Variables” a list and description of the variables included in the maternal database is included and, in the sheet, named “Neonatal Variables” a list and description of the variables included in the neonatal database is included.</p> <p>Three files of “R” extension and one “rmd” are included. The file named “RunningModelsFunctions.R” is the one use to run the models that are included in the analyses, the file named “2. Final Analysis LCG Trial.R” is the one in which the tables are prepared, and the file named “3. LCG Results Output Final.rmd” is used to export the tables with results. The R file named “funciones.tablas.R” is used in the analyses.</p>
Data from: A novel approach to quantifying mammal locomotor repertoires using scoring and cluster analysis
Open the record for dataset details and reuse information.
Extended data tables to Haering and Habermann, F1000Res, RNfuzzyApp: an R shiny RNA-seq data analysis app for visualisation, differential expression analysis, time-series clustering and enrichment analysis
Open the record for dataset details and reuse information.
Characteristic and spatiotemporal variation of air pollution in Northern China based on correlation analysis and clustering analysis of five air pollutants
<p>Data for "Characteristic and spatiotemporal variation of air pollution in Northern China based on correlation analysis and clustering analysis of five air pollutants"</p>
A K-means Clustering Analysis of the Jovian and Terrestrial Magnetopauses
<p>Jovian magnetopause crossings used in the referenced paper and a README file</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.