Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,326
datasets available to search
ShareScore release 0.9.0
Dataset results
2,326 results for “clusters”
FIGURE 51. Bayesian clustering for K in A taxonomic revision of the Palaearctic species of the ant genus Tapinoma Mayr 1861 (Hymenoptera: Formicidae)
FIGURE 51. Bayesian clustering for K=2 of 15 microsatellite loci of Tapinoma madeirense (orange) and T. subboreale (blue).
A Comparison Between Hierarchical and Non-Hierarchical Software Clustering
<p>Dataset for "A Comparison Between Hierarchical and Non-Hierarchical Software Clustering" paper results.</p>
Supplementary Material - A Comparison of Clustering Approaches for Metaheuristic Behaviour Data
<p>This repo contains all results data as images and tables that was gathered and used for evaluating the experiments in the paper "A Comparison of Clustering Approaches for Metaheuristic Behaviour Data". The folder structure is representative of the respective dataset used to generate the results.</p>
Blueprint for modelling time effects in a stepped-wedge, cluster-randomised trial
<div><strong># Project title</strong></div> <div>Blueprint for modelling time effects in a stepped-wedge, cluster-randomised trial</div> <div> </div> <div><strong># Brief project overview</strong></div> <div>This is Data coming out of the IPOS-Dem project, a Stepped-wedge Cluster Randomised Trial (SW-CRT). Please make shure to check out the SW-CRT protocol for more details at <a href="https://www.doi.org/10.1111/jan.14953">https://www.doi.org/10.1111/jan.14953</a>. </div> <div> </div> <div>We are currently working on our main study (SW-CRT) analysis. While we remain dedicated to open science, our priority is to ensure the utmost integrity of our ongoing research.</div> <div> </div> <div><strong># Contributors</strong></div> <div>Frank Spichiger, Andrea Koppitz, André Meichtry</div> <div> </div> <div><strong>## Repository overview</strong></div> <div> </div> <div>|-- readme.md</div> <div>|-- documentation</div> <div> |--readme.md</div> <div> |--Datadictionary.html</div> <div>|-- data</div> <div> |--readme.md</div> <div> |--IPOS-Dem_CH-SW-CRT_PLD_Qualidem.feather</div> <div> |--IPOS-Dem_CH-SW-CRT_PLD_locf_Qualidem.feather</div> <div> |--csv</div> <div> |--IPOS-Dem_CH-SW-CRT_PLD_Qualidem.csv</div> <div> |--IPOS-Dem_CH-SW-CRT_PLD_locf_Qualidem.csv</div> <div>|-- analysis</div> <div> |--readme.md</div> <div> |--240229_Analysis_Main.html</div> <div> |--240229_Analysis_Main.qmd</div> <div> |--grateful-refs.bib</div> <div> |--output</div> <div> |--models-REML.csv</div> <div> |--models.csv</div> <div> |--figures</div> <div> |--Conversions for plotting - 1.png</div> <div> |--Define Model 5-1.png</div> <div> |--Defining Model 4-1.png</div> <div> |--m1predictions.eps</div> <div> |--m1predictions.png</div> <div> |--m2predictions.eps</div> <div> |--m2predictions.png</div> <div> |--m3predictions.eps</div> <div> |--m3predictions.png</div> <div> |--m4predictions.eps</div> <div> |--m4predictions.png</div> <div> |--m5predictions.eps</div> <div> |--m5predictions.png</div> <div> |--Model 1 plot-1.png</div> <div> |--Model 1 plot-2.png</div> <div> |--Model 2 plot-1.png</div> <div> |--Model 2 plot-2.png</div> <div> |--Model 3 plot-1.png</div> <div> |--More plots for Model 1-1.png</div> <div> |--More plots for Model 1-2.png</div> <div> |--More plots for Model 4-1.png</div> <div> |--More plots for Model 5-1.png</div> <div> |--Plots for Model 4-1.png</div> <div> |--Plots for Model 5-1.png</div> <div> |--Plots for the other two models-1.png</div> <div> |--Plots for the other two models-2.png</div> <div> </div> <div> </div> <div># Applicable instructions</div> <div>The R markdown files were generated using R 4.2.2 in RStudio 2023.12.1 for MacOS X you can run them using free software but will need to install R, Rstudio and the packages cited at the end of the quarto or the rendered html we used.</div> <div> </div> <div># Additional resources</div> <div>- Main study protocol: <a href="https://www.doi.org/10.1111/jan.14953">https://www.doi.org/10.1111/jan.14953</a></div> <div>- QUALIDEM measure used: <a href="https://doi.org/10.1186/1477-7525-11-91">https://doi.org/10.1186/1477-7525-11-91</a> </div> <div>- Pos-Pal consortium with more information on the measure: <a href="https://www.pos-pal.org">https://www.pos-pal.org</a></div> <div>- IPOS-Dem translation and adaption: <a href="https://www.doi.org/10.1186/s41687-022-00420-7">https://www.doi.org/10.1186/s41687-022-00420-7</a></div> <div>- IPOS-Dem inter-rating reliability: <a href="https://doi.org/10.1371/journal.pone.0286557">https://doi.org/10.1371/journal.pone.0286557 </a></div>
Simulation Results in the Paper "Propagation of Slow Slip Events on Rough Faults: Clustering, Back Propagation, and Re-rupturing" [Dataset]
<p>Data file "simulations.mat" contains the 5 simulations of slow slip events on flat or rough faults. </p> <table> <tbody> <tr> <td>structure array</td> <td>description</td> <td>reference</td> </tr> <tr> <td>s0</td> <td> 2.5 km long flat fault</td> <td>Fig. 2b</td> </tr> <tr> <td>s1</td> <td>2.5 km long rough fault</td> <td>Fig. 2c</td> </tr> <tr> <td>s2</td> <td>10 km long rough fault</td> <td>Fig. 5</td> </tr> <tr> <td>s3</td> <td>10 km long fractal fault</td> <td>Fig. 7</td> </tr> </tbody> </table> <p>structure array consists of:</p> <p>t: time (s)</p> <p>x: distance (m)</p> <p>v: slip rate (m/s)</p> <p>slip: accumulated slip (m)</p> <p>tau: shear stress (Pa)</p> <p>sigma: normal stress (Pa)</p> <p>notes: description</p> <p> </p> <p>Data file "catalog.mat" contains 3 simulated slow slip events' catalogs on flat and rough faults, c0, c1, and c2, in Fig. 4a, 4b, and 4c, respectively.</p> <p>It consists of:</p> <p>l: rupture length (m)</p> <p>time: time (s)</p> <p>notes: description</p>
Electron number density, temperature, gas mass and hydrostatic mass profiles for 22 eROSITA clusters
Open the record for dataset details and reuse information.
Seminar - Dr Thomas Beaney - Simplifying complexity: generating multi-resolution clusters of chronic diseases in 10 million people in England
<p>Recording of the presentation given on Wednesday 23rd October by Dr <span>Thomas</span> <span>Beaney</span> “Simplifying compl<span>e</span>xity: g<span>e</span>n<span>e</span>rating multi-r<span>e</span>solution clust<span>e</span>rs of chronic dis<span>e</span>as<span>e</span>s in 10 million p<span>e</span>opl<span>e</span> in <span>E</span>ngland”<br><br>In this talk, Tom discusses an approach for g<span>e</span>n<span>e</span>rating clust<span>e</span>rs of chronic dis<span>e</span>as<span>e</span>s in p<span>e</span>opl<span>e</span> with Multipl<span>e</span> Long-T<span>e</span>rm Conditions (MLTC), which combin<span>e</span>s natural languag<span>e</span> proc<span>e</span>ssing and n<span>e</span>twork-bas<span>e</span>d clust<span>e</span>ring. H<span>e</span> also discusses th<span>e</span> chall<span>e</span>ng<span>e</span>s of <span>e</span>valuating and using clust<span>e</span>rs in practic<span>e</span> to influ<span>e</span>nc<span>e</span>pati<span>e</span>nt car<span>e</span>.</p> <p>You can view the publication of his work here: https://www.nature.com/articles/s43856-024-00529-4</p>
Pangenomes of multiple species for the "Cluster efficient pangenome graph construction with nf-core/pangenome" manuscript.
<p>Pangenomes of multiple species for the "Cluster efficient pangenome graph construction with nf-core/pangenome" manuscript.</p> <p>Each pangenome is represented in a FASTA format file. Each FASTA file was compressed with <em>bgzip</em> and subsequent indices were created with <em>tabix</em>.</p> <p>The name of each file specifies:</p> <ul> <li>The species,</li> <li>and the number of haplotypes.</li> </ul> <p>The built pangenomes graph are represented in GFA format. Each pangenome graph in the paper is uploaded here, too.</p>
Data for "The Massive and Distant Clusters of WISE Survey. XII. Exploring X-ray AGN in Dynamically Active Massive Galaxy Clusters at z~1"
<p>This dataset corresponds to the manuscript titled "The Massive and Distant Clusters of <em>WISE</em> Survey. XII. Exploring X-ray AGN in Dynamically Active Massive Galaxy Clusters at z~1" (<span>DOI: 10.3847/1538-4357/adbae4)</span>, accepted for publication in <em>The Astrophysical Journal</em> on February 24, 2025. To replicate the plots and access the catalogs used in the paper, download and extract all the zip folders along with the Jupyter Notebook "<a href="https://zenodo.org/api/records/14928038/draft/files/madcows_notebook.ipynb/content">madcows_notebook.ipynb</a>" into the same directory. Open the notebook and follow the instructions provided. For any issues, please contact the corresponding author.</p>
Data Collaborative and Clusters_Bartolomucci_Bresolin
<p> The dataset describes 171 Data Collaboratives according to organisational, technological, geographical, and impact dimensions. The dataset is derived from the one present on datacollaboratives.org and complemented with additional data obtained by secondary data sources. </p> <p>A related work using a previous version of this dataset can be found online with the following title. Fostering data collaboratives’ systematisation through models’ definition and research priorities setting." <em>DG. O 2022: The 23rd Annual International Conference on Digital Government Research</em>. 2022</p> <p> </p>
Three-Dimensional Segmentation Assisted with Clustering Analysis for Surface and Volume Measurements of Equine Incisor in Multidetector Computed Tomography Data Sets
<p>The dataset contains computed tomography (CT) images of head horses with annotations of 12 segments corresponding to areas with teeth. Imaged animals: 49 horses. Measured animal features such as surface area, and volume are included. The study was supported by the National Science Centre, Poland as a part of the project<br>Miniatura 6 No 2022/06/X/ST6/00431. </p>
Global Landside Clustering of Aquaculture Ponds Distribution Acquired from Dense Time-Series Sentinel-2 Images by Google Earth Engine
<p>This dataset reveals the global distribution pattern of landside clustering aquaculture ponds (LCAP) from a spatial perspective for the first time. It was derived from 4,015,054 tiles of the 10-m Sentinel-2 time-series images collected throughout 2020. The total area of global LCAP was estimated at 55,337.03 km2. Accuracy verification revealed that the Omission Error and Commission Error of the data is 7.51% and 16.69% respectively. We provide this dataset in <em>ESRI</em> <em>shapefile </em>format (.zip), which can be opened by <em>ArcGIS. </em>We invite you to download and utilize this dataset and recommend citing the following two references.</p>
Explainable Clustering Applied to the Definition of Terrestrial Biomes - data
<p>Data used for analysis in "Explainable Clustering Applied to the Definition of Terrestrial Biome" - using Decision Tree and Clustering techniques to identify biomes.</p> <p>Land surface properties:</p> <ul> <li><strong>TreeCover </strong>- Vegetation Continuous Fields (VCF) collection 6 fractional tree cover from DiMiceli et al. 2015, regridded as per Kelley et al. 2019</li> <li><strong>NonTreeCover </strong>- VCF fractional herb cover</li> <li><strong>Urban </strong>cover from the History Database of the Global Environment, Version 3.1 (HYDE) Klein Goldewijk et al. 2011</li> <li><strong>Crop </strong>cover (from HYDE)</li> <li><strong>Pasture </strong>Cover (from HYDE)</li> <li><strong>PopDen </strong>(population density from HYDE)</li> </ul> <p>Climate:</p> <ul> <li><strong>MAP_CRU </strong>- Mean annual precipitation from version 4.01 of the Climatic Research Unit Time Series high resolution gridded dataset (CRU TS v4.01) (Harris & Jones 2017)</li> <li><strong>MAT </strong>- Mean annual temperature from CRU)</li> <li><strong>MADD_CRU </strong>- Mean annual dry days from CRU - i.e seasonality of rainfall</li> <li><strong>MTWM </strong>- Mean Maximum Temperature of the warmest month from CRU</li> <li><strong>MTCM </strong>- Mean Minumum Temperature of the coldest month from CRU</li> <li><strong>SW1 </strong>- direct downwards SW simulated using the SLASH model using CRU cload cover</li> <li><strong>SW2 </strong>- diffuse downwards SW simulated using the SLASH model using CRU cload cover</li> <li><strong>BurntArea_GFED_four_s </strong>- Burnt area from Global Fire Emissions Database, Version 4.1 (GFEDv4.1) (Van Der Werf et al. 2017)</li> <li><strong>MaxWind</strong><em><strong> </strong></em>(Mean Max Windspeed from CRU-(National Centers for Environmental Prediction (Harris 2019)</li> </ul> <p>Dimiceli, C., Carroll, M., Sohlberg, R., Kim, D. H.,Kelly, M., and Townshend, J. R. G. (2015). Mod44bmodis/terra vegetation continuous fields yearly l3global 250m sin grid v006 (v006).</p> <p>Harris, I. (2019). CRU JRA v1. 1: A forcings dataset ofgridded land surface blend of Climatic Research Unit (CRU) and Japanese reanalysis (JRA) data, January1901–December 2017, University of East Anglia Climatic Research Unit, Centre for Environmental DataAnalysis.</p> <p>Harris, I. and Jones, P. (2017). CRU TS4. 01: Climatic Re-search Unit (CRU) Time-Series (TS) version 4.01 ofhigh-resolution gridded data of month-by-month vari-ation in climate (Jan. 1901–Dec. 2016).Centre forEnvironmental Data Analysis, 25.</p> <p>Kelley, D. I., Bistinas, I., Whitley, R., Burton, C., Marthews,T. R., and Dong, N. (2019). How contemporary biocli-matic and human controls change global fire regimes</p> <p>Klein Goldewijk, K., Beusen, A., Van Drecht, G., and DeVos, M. (2011). The HYDE 3.1 spatially explicitdatabase of human-induced global land-use changeover the past 12,000 years.Global Ecology and Bio-geography, 20(1):73–86.</p> <p>Van Der Werf, G. R., Randerson, J. T., Giglio, L.,Van Leeuwen, T. T., Chen, Y., Rogers, B. M., Mu, M.,Van Marle, M. J., Morton, D. C., Collatz, G. J., et al.(2017). Global fire emissions estimates during 1997–2016.Earth System Science Data, 9(2):697–720.</p>
Vertical profiles of 3D cluster properties
<p>The two zip files contain the properties of the 3D clusters used in the submission version of the Paper "Size-dependence of surface-rooted three-dimensional convective objects in continental shallow cumulus simulations". The paper was submitted to JAMES in May 2021.</p> <p>When unzipped, the data is stored in python panda dataframes in pkl files. Warning! Unzipping the files increases the data size by a factor of 20.</p> <p>The two python scripts contain the functions used to segment the 3D snapshots into individual objects. The most interesting things are in proc_watershed, cusize_functions is just a collection of mostly abandoned functions. The functions in proc_watershed are reasonably well commented.</p>
New crater clusters on Mars
<p>Crater cluster data set used in the submitted paper</p>
Finite Mixture Models for Clustering Sales Series Data in the Presence of Promotions
<pre> 131 sales data where promotion causes volatility over the entire time series</pre>
Birth cluster simulations of planetary systems with multiple super-Earths: initial conditions for white dwarf pollution drivers
<p>We provide the output parameters from our planetary system simulations around white dwarf main-sequence progenitors (1.5, 2.0 and 2.5 solar mass stars) embedded in a birth star cluster containing 8,000 stars. The data can serve the community as initial conditions for subsequent evolution simulations, e.g. to further investigate the role of eccentric planets in the pollution of white dwarfs. The planetary systems, consisting solely of super-Earths with 0.01 Jovian masses, were started in three different orbital configurations. The 3P model contains 3 planets with initial semimajor axes between 2.00 and 18.63 AU. We also simulated two 7-planet system models with different orbital spacing. The 7PC model represents a compact planetary system with 7 planets between 2.00 and 18.63 AU. The 7PW model represents a wide planetary system with 7 planets between 2.00 and 56.87 AU. The three planetary system models were distributed around 193 stars with 1.5 solar masses, around 114 stars with 2.0 solar masses and around 101 stars with 2.5 solar masses, and numerically integrated for 100 Myr considering the gravitational forces from the neighbouring stars in the star cluster.</p> <p>See the publication Stock et al. (2022) for further information.</p>
Effectiveness of a multicomponent intervention consisting of education and feedback on reducing benzodiazepine prescriptions by general practitioners: BENZORED hybrid type I cluster randomized controlled trial.
<p>Complete dataset variables:</p> <p> </p> <p>GP_ID<br> Health_District<br> Health_District_name<br> PHC_ID<br> PHC_ID_name<br> Arm<br> DHD_Baseline<br> DHD_12m<br> PercentageBZD_baseline<br> PercentageBZD_12m<br> PercentageBZD_baseline_age65<br> PercentageBZD_12m_age65</p>
FIGURE. Scatter plots (N=200) and linear regression lines of the length and diameter of termite coprolites from the Lower Cretaceous Huolinhe Formation in eastern Inner Mongolia, China. The grey shading represents the 95% confidence interval of linear relationship. Note scatter plots depicting a k-means clustering analysis reveals three groups, indicated by circles of different colours; stars of different colour mean the clusters centroids which are the average length and diameter. in Termite coprolites (Blattodea: Isoptera) from the Early Cretaceous of eastern Inner Mongolia, Northeast China
FIGURE. Scatter plots (N=200) and linear regression lines of the length and diameter of termite coprolites from the Lower Cretaceous Huolinhe Formation in eastern Inner Mongolia, China. The grey shading represents the 95% confidence interval of linear relationship. Note scatter plots depicting a k-means clustering analysis reveals three groups, indicated by circles of different colours; stars of different colour mean the clusters centroids which are the average length and diameter.
S1_The_differentially_expressed_genes_enriched_in_GO_clusters
<p>This is supporting information to the article <em>Downregulation of ammonium uptake improves the growth and tolerance of Kluyveromyces marxianus at high temperature</em>, which has been submitted to <em>MicrobiologyOpen</em>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.