Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “k-means clustering”
Integrated freshwater abundance and connectivity clusters at the Hydrologic Unit 8 scale for the Midwest and Northeast U.S.A. – freshwater metric variables and k-means cluster assignment
This dataset includes integrated freshwater abundance and connectivity cluster output, principal component scores, and lake, wetland, and stream abundance and connectivity metrics measured at the Hydrologic Unit 8 (HU8) scale for 17 U.S. states in the Midwest and Northeast regions (appr. 1,800,000 km2). The intent of the cluster analysis is to characterize the macroscale patterns of the integrated freshwater landscape that includes lakes, wetlands, and streams and their surface connectivity attributes. We define freshwater connectivity as the permanent surface hydrologic connections that link lakes, wetlands, and streams and measure connectivity as the landscape position of systems within stream networks. Geographic data used in the analysis are in LAGOS-NE-GEO database v. 1.03 (Lake multi-scaled geospatial and temporal database), an integrated, multi-thematic geographic database (Soranno et al. 2015). The integrated freshwater clusters were created through a multi-step process as follows: 1) we quantified multiple freshwater connectivity metrics for lakes, streams, and wetlands separately, 2) we performed principal components analysis (PCA) on the connectivity metric values for each freshwater type to reduce collinearity, and 3) we performed k-means cluster analysis to group spatial units with similar freshwater connectivity characteristics. The resulting freshwater clusters are representations of the macroscale patterns of freshwater abundance and connectivity in the landscape.
Freshwater connectivity clusters for lakes, wetlands, and streams at the Hydrologic Unit 12 scale in the Midwest and Northeast U.S.A. – freshwater metric variables and K-means cluster assignment
This dataset includes freshwater connectivity cluster output and principal component scores for lakes, wetlands, and streams measured at the Hydrologic Unit 12 (HU12) scale in 17 U.S. states in the Midwest and Northeast regions (appr. 1,800,000 km2). The intent of the cluster analysis is to characterize the macroscale patterns of freshwater connectivity attributes. We define freshwater connectivity as the permanent surface hydrologic connections that link lakes, wetlands, and streams and measure connectivity as the landscape position of systems within stream networks. Geographic data used in the analysis are in LAGOS-NE-GEO database v. 1.03 (Lake multi-scaled geospatial and temporal database), an integrated, multi-thematic geographic database (Soranno et al. 2015). Freshwater connectivity clusters were created separately for lakes, wetlands, and streams through a multi-step process as follows: 1) we quantified multiple freshwater connectivity metrics, 2) we performed principal components analysis (PCA) on the connectivity metric values for each freshwater type to reduce collinearity, and 3) we performed k-means cluster analysis to group spatial units with similar freshwater connectivity characteristics. The resulting freshwater clusters are representations of the macroscale patterns of lake, wetland, and stream connectivity in the landscape.
ACCESS-AM2 Southern Ocean cloud and radiation data for k-means clustering and analysis
<p>The ACCESS-AM2 (Australian Community Climate and Earth-System Simulator - Atmospheric Model Version 2) data and k-means analysis used for the study described in Fiddes et al. 2022 '<em>Southern Ocean cloud and shortwave radiation biases in a nudged climate model simulation: does the model ever get it right?' .</em> </p> <p>Included files: </p> <ul> <li>modis_cluster_centres_2015-2019.nc - kmeans derived cluster centres for MODIS</li> <li>modis_cluster_labels_2015-2019.nc - kmeans derived cluster labels for MODIS </li> <li>bx400_cluster_labels_2015-2019.nc - kmeans fitted cluster label for model </li> <li>COSP_vars_bx400_2015-2019.nc - model data for analysis </li> </ul> <p>The code that performs the analysis/generates this data and has instructions for where to download MODIS data can be found here: https://github.com/sfiddes/code_for_publications_2022/tree/main/ACCESS_cloud_radiation_eval</p>
A K-means Clustering Analysis of the Jovian and Terrestrial Magnetopauses
<p>Jovian magnetopause crossings used in the referenced paper and a README file</p>
K-Means Clustering Peraturan Kementerian
<p>Analisis Peraturan Kementerian menggunakan metode K-Means Clustering</p>
FIGURE. Scatter plots (N=200) and linear regression lines of the length and diameter of termite coprolites from the Lower Cretaceous Huolinhe Formation in eastern Inner Mongolia, China. The grey shading represents the 95% confidence interval of linear relationship. Note scatter plots depicting a k-means clustering analysis reveals three groups, indicated by circles of different colours; stars of different colour mean the clusters centroids which are the average length and diameter. in Termite coprolites (Blattodea: Isoptera) from the Early Cretaceous of eastern Inner Mongolia, Northeast China
FIGURE. Scatter plots (N=200) and linear regression lines of the length and diameter of termite coprolites from the Lower Cretaceous Huolinhe Formation in eastern Inner Mongolia, China. The grey shading represents the 95% confidence interval of linear relationship. Note scatter plots depicting a k-means clustering analysis reveals three groups, indicated by circles of different colours; stars of different colour mean the clusters centroids which are the average length and diameter.
K-means clustering of the DEMIX data set
The DEMIX database (Guth, 2023) contains statistics from 6 test 1 arc second DEMs (ALOS, ASTER, CopDEM, FABDEM, NASADEM, and SRTM) compared to high resolution reference DEMs. The database contains 236 DEMIX tiles (Guth and others, 2023) and has 7 tile characteristics for each tile. A K-means clustering of the database, and an additional set of 4 land cover and landform classifications for the 236 DEMIX tiles produced variable numbers of clusters. This data set includes five databases including the clustering results, a figure showing the scatterplots showing the relations among the tile characteristics in the database, and map showing the location of the clusters. References: Guth, P. L., 2023. DEMIX GIS Database Version 2 (2.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.8062008 Guth, Peter L., Peter Strobl, Kevin Gross, & Serge Riazanoff. (2023). DEMIX 10k Tile Data Set (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7504791
Dataset and code for "FK-means: Automatic Atrial Fibrosis Segmentation using Fractal-guided K-means Clustering with Voronoi-Clipping Feature Extraction of Anatomical Structures": FKmeans for fibrosis segmentation
<p>Assessment of left atrial (LA) fibrosis from late gadolinium enhancement (LGE) magnetic resonance imaging (MRI) adds to the management of patients with atrial fibrillation (AF). However, accurate assessment of fibrosis in the LA wall remains challenging. Excluding anatomical structures in the LA proximity using clipping techniques can reduce misclassification of LA fibrosis. A novel FK-means approach for combined automatic clipping and automatic fibrosis segmentation was developed. This approach combines a feature-based Voronoi diagram with a hierarchical 3D K-means fractal-based method. The proposed automatic Voronoi clipping method was applied on LGE MRI data and achieved a Dice score of 0.75, similar as the score obtained by a deep learning method (3D UNet) for clipping (0.74). The automatic fibrosis segmentation method, which utilizes the Voronoi clipping method, achieved a Dice score of 0.76. This outperformed a 3D U-Net method for clipping and fibrosis classification, which had a Dice score of 0.69. Moreover, the proposed automatic fibrosis segmentation method achieved a Dice score of 0.90, using manual clipping of anatomical structures. The findings suggest that the automatic FK-means analysis approach enables reliable LA fibrosis segmentation and that clipping of anatomical structures in the atrial proximity can add to the assessment of atrial fibrosis. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.