Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,326

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,326 results for “clusters”

Learn how ShareScore rates datasets ↗
zenodo36/100

A possible formation channel for blue hook stars in globular cluster - II. Effects of metallicity, mass ratio, tidal enhancement efficiency and helium abundance

<p>MESA inlists and run_star_extras associated with <a href="https://ui.adsabs.harvard.edu/?#abs/2016MNRAS.463.3449L">Lei et al. (2016)</a>. MESA version 7211.</p> <p>Publication DOI:&nbsp;<a href="https://doi.org/10.1093/mnras/stw2242">10.1093/mnras/stw2242</a></p>

opencc-by-4.0Mar 2019View details →
zenodo36/100

Data for manuscript: The orbital anisotropy profiles of nearby globular clusters from Gaia Data Release 2

<p>We upload the data used in our paper here so that our results may be reproduced. We include the dataset of stars that survive our cuts, the profiles we plot, and the manual points selected as part of our CMD cut. See the paper for details. The first version of this paper is published on the arXiv with ID:&nbsp;arXiv:1903.11070.&nbsp;</p>

opencc-by-4.0May 2019View details →
zenodo36/100

Performance Analysis of Single Board Computer Clusters

<p>This dataset contains the outputs from HPL runs used to measure performance of 16 nodes clusters built using Raspberry Pi 3 Model B, Raspberry Pi 3 Model B+, and Odroid C2.&nbsp; These clusters were constructed using the Pi Stack PCB which is available from https://doi.org/10.5258/SOTON/D0379.</p>

opencc-by-4.0Dec 2018View details →
zenodo36/100

Text-fig. 2 Cluster diagram of the considered acritarch assemblages. Paired group, Jaccard measure. in Overview Of The Stratigraphy And Initial Quantitative Biogeographical Results From The Devonian Of The Albergaria-A-Velha Unit (Ossa-Morena Zone, W Portugal)

Text-fig. 2 Cluster diagram of the considered acritarch assemblages. Paired group, Jaccard measure.

opencc-by-4.0Dec 2008View details →
zenodo36/100

MMSEQS meets AntiRef90: reference clusters of human antibody sequences

<p>This data set contains pre-computed mmseqs databases for the antiref fasta files created by <em>Briney et al.</em>&nbsp;</p> <p>Please cite the original work if you use any of the databases provided here.</p> <p>Sources:</p> <ul> <li><a href="https://github.com/brineylab/antiref">Antiref GitHub</a></li> <li><a href="../records/7474336">Antiref Zenodo</a></li> <li><a href="https://academic.oup.com/bioinformaticsadvances/article/3/1/vbad109/7247530?login=true">Antiref Paper</a></li> </ul> <p>&nbsp;</p> <p>The mmseqs databases were created as follows:</p> <p>&nbsp;</p> <p>```</p> <p>aria2x -x16 -s16 --input-file antiref_links.txt<br>snakemake -s antiref_mmseqs.smk --jobs 1 --cores 1 --local-cores 250</p> <p>```</p> <p>&nbsp;</p> <p>Please check the summary repository for the fasta files and snakemake files. In this sub repo we only store the antiref files matching the title of the repo.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

ELF3 prion-like domain Martini clustering simulations

<p>This dataset contains Martini coarse-grain simulations of ELF3-PrD, with each simulation containing 100 PrD monomers. These trajectories were created as part of a publication exploring the temperature-responsive condensation of the ELF3-PrD in the scientific pulication titled, "Molecular dynamics simulations illuminate the role of sequence context in the ELF3-PrD-based temperature sensing mechanism in plants." Included are trajectories for ELF3-PrD variants including wildtype (7 glutamine-long polyQ tract), 0Q (variable poly-glutamine tract removed), 19Q (polyQ tract extenden to 19 glutamine residues) and the F527A mutant. Each variant includes trajectories at 290K, 300K, 320K and 405K. There are three replicates for each condition, except for wildtype 300K and 19Q 340K, of which two replicates are provided.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Production of Alternate Realizations of DESI Fiber Assignment for Unbiased Clustering Measurement in Data and Simulations

<p>A critical requirement of spectroscopic large scale structure analyses is correcting for selection of which galaxies to observe from an isotropic target list. This selection is often limited by the hardware used to perform the survey which will impose angular constraints of simultaneously observable targets, requiring multiple passes to observe all of them. In SDSS this manifested solely as the collision of physical fibers and plugs placed in plates. In DESI, there is the additional constraint of the robotic positioner which controls each fiber being limited to a finite patrol radius. A number of approximate methods have previously been proposed to correct the galaxy clustering statistics for these effects, but these generally fail on small scales. &nbsp;To accurately correct the clustering we need to upweight pairs of galaxies based on the inverse probability that those pairs would be observed (Bianchi &amp; Percival 2017). This paper details an implementation of that method to correct the Dark Energy Spectroscopic Instrument (DESI) survey for incompleteness. To calculate the required probabilities, we need a set of alternate realizations of DESI where we vary the relative priority of otherwise identical targets. &nbsp;These realizations take the form of alternate Merged Target Ledgers (AMTL), the files that link DESI observations and targets. We present the method used to generate these alternate realizations and how they are tracked forward in time using the real observational record and hardware status, propagating the survey as though the alternate orderings had been adopted. We detail the first applications of this method to the DESI One-Percent Survey (SV3) and the DESI year 1 data. We include evaluations of the pipeline outputs, estimation of survey completeness from this and other methods, and validation of the method using mock galaxy catalogs.&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

InZePro Cluster InForm Referencedataset

<p>Referencedataset of the BMBF Cluster InZePro of the project InForm - <strong>In</strong>telligent<strong> Form</strong>ation system<strong> </strong></p> <p>The dataset contains data of the formation process of different cells types (coin cell, pat-cell, pouch cell, prismatic cell PHEV1).</p> <p>On a subset a special End-of-Line test directly after formation.</p> <p>Further details are provided in the excel file.</p> <p>Raw data is compressed in the zip-File.</p> <p>The hdf5-file combines all cell data and can be opened i.e:<br> import hdfdict</p> <p>hdf5name = &#39;&#39; ReferenceDatasetInForm.h5 &quot;</p> <p>hdf5dict = dict(hdfdict.load(hdf5name, mode=&#39;r+&#39;,lazy=False))</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Support data for conference paper "Service-Oriented Model for Handling mMTC Subscribers' Traffic in a 5G Cluster"

<p>Support data for conference paper<br>V. Kovtun, and O. Kovtun, &ldquo;Service-Oriented Model for Handling mMTC Subscribers&rsquo; Traffic in a 5G Cluster.&rdquo; In Proc. 5th 5th International Workshop on Intelligent Information Technologies &amp; Systems of Information Security, CEUR-WS, vol. 3675, 2024; pp. 236-246.<br>This research is part of the project No. 2022/45/P/ST7/03450 co-funded by the National Science Centre and the European Union Framework Programme for Research and Innovation Horizon 2020 under the Marie Skłodowska-Curie grant agreement No. 945339.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Support data for article "The concept of network resource control of a 5G cluster focused on the smart city's critical infrastructure needs"

<p>Support data for article:</p> <div>V. Kovtun, K. Grochla, and K. Połys, &ldquo;The concept of network resource control of a 5G cluster focused on the smart city&rsquo;s critical infrastructure needs,&rdquo; Alexandria Engineering Journal, vol. 94. Elsevier BV, pp. 248&ndash;256, May 2024. doi: 10.1016/j.aej.2024.03.038.</div> <p>This research is part of the project No. 2022/45/P/ST7/03450 co-funded by the National Science Centre and the European Union Framework Programme for Research and Innovation Horizon 2020 under the Marie Skłodowska-Curie grant agreement No. 945339.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Support data for conference paper "The Concept of Efficient Utilization of the Uplink Frequency Resource of a Smart Factory 5G Cluster by IIoT Devices"

<p>upport data for conference paper<br>Kovtun, and O. Kovtun, &ldquo;The Concept of Efficient Utilization of the Uplink Frequency Resource of a Smart Factory 5G Cluster by IIoT Devices.&rdquo; In Proc. 8th International Conference on Computational Linguistics and Intelligent Systems. Volume I: Machine Learning Workshop, CEUR-WS, vol. 3664, 2024; pp. 273-283.<br>This research is part of the project No. 2022/45/P/ST7/03450 co-funded by the National Science Centre and the European Union Framework Programme for Research and Innovation Horizon 2020 under the Marie Skłodowska-Curie grant agreement No. 945339.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Cluster merged fluxgate/search coil data for a solar wind interval occuring on 15/02/2015 21:25:00-22:40:00

<p>Cluster merged fluxgate/search coil data for a solar wind interval occuring on 15/02/2015 21:25:00-22:40:00</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Efficient parameterization of transferable Atomic Cluster Expansion for water

<p>This collection contains files associated with &nbsp;Journal of Chemical Theory and Computation. "Efficient parameterization of transferable Atomic Cluster Expansion for water" (2024) paper:</p> <p>- ACE potentials for water.</p> <p>-Active set inverted (ASI) for the ACE potential</p> <p>- Training dataset that was used for fitting Atomic Cluster Expansion potential.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Efficient Enumeration of the Optimal Solutions to the Correlation Clustering problem

<div><strong>Description. </strong>This is the data used in the experiments presented in the following paper:<br> <div> <ul> <li> <div> <div>N. Arınık, R. Figueiredo, and V. Labatut, &ldquo;Efficient Enumeration of the Optimal Solutions to the Correlation Clustering problem,&rdquo; <em>Journal of Global Optimization</em>, vol. 86, pp. 355&ndash;391, 2023. DOI: <a href="http://doi.org/10.1007/s10898-023-01270-3">10.1007/s10898-023-01270-3</a> ⟨<a href="https://hal.science/hal-03935831">hal-03935831</a>⟩</div> </div> </li> </ul> <p><strong>Source code. </strong>The related source code is available on GitHub:&nbsp;</p> <ul> <li><a href="https://figshare.com/articles/dataset/Efficient_Enumeration_of_Correlation_Clustering_Optimal_Solution_Space/%3Ci%3Ehttps://github.com/%3C/i%3E%3Ci%3E%3Ci%3ECompNet%3C/i%3E/Sosocc%3C/i%3E">https://github.com/CompNet/Sosocc</a></li> <li><a href="https://figshare.com/articles/dataset/Efficient_Enumeration_of_Correlation_Clustering_Optimal_Solution_Space/%3Ci%3Ehttps://github.com/CompNet/EnumCC%3C/i%3E">https://github.com/CompNet/EnumCC</a></li> </ul> <div><strong>Citation. </strong>If you use these data, please cite the above reference:</div> <div><br><code>@Article{Arinik2021,</code><br><code>&nbsp; author &nbsp; &nbsp;= {Arınık, Nejat and Figueiredo, Rosa and Labatut, Vincent},</code><br><code>&nbsp; title &nbsp; &nbsp; = {Efficient Enumeration of the Optimal Solutions to the Correlation Clustering problem},</code><br><code>&nbsp; journal &nbsp; = {Journal of Global Optimization},</code><br><code>&nbsp; year &nbsp; &nbsp; &nbsp;= {2023},</code><br><code>&nbsp; volume &nbsp; &nbsp;= {86},</code><br><code>&nbsp; pages &nbsp; &nbsp; = {355-391},</code><br><code>&nbsp; doi &nbsp; &nbsp; &nbsp; = {10.1007/s10898-023-01270-3},</code><br><code>}</code></div> <div>&nbsp;</div> <div><strong>Funding. </strong>This research benefited from the support of Agorantic FR 3621, as well as the FMJH Program PGMO and from the support to this program from EDF-THALES-ORANGE-CRITEO.</div> <div>&nbsp;</div> <div><strong>Further details. </strong>We describe below the structure of the zip file `article_materials.zip`:</div> <div> <ul> <li>Experiments for Dataset 1 <ul> <li>delay_exec_time: all the results and plots regarding the difference of execution times between EnumCC(3) and OneTreeCC() (i.e., EnumCC(3) minus OneTreeCC()), represented on the log-scaled y-axis of the plots. When such difference takes a negative value, this means our proposed method EnumCC(3) runs faster than OneTreeCC().</li> <li>EnumCC_nb-jumps: all the results regarding the number of jumps related to EnumCC(3), i.e. n<sub><em>jump</em></sub>(EnumCC(3))</li> <li>exec_time: all the results regarding the execution times of EnumCC(3) and OneTreeCC().</li> <li>nb-sols: all the results regarding the number of optimals solutions based on EnumCC(3). Note that we show the results of OneTreeCC() only for those with n=50, since both methods run out of the time limit of 12h for several networks with n=50.</li> </ul> </li> <ul> <li>networks: We generate these complete and incomplete networks through our random signed network generator, which is publicly available online. For complete unweighted signed networks, this model relies on only three parameters: n (number of vertices), l<sub>0</sub>&nbsp;(initial number of modules) and&nbsp;<em>q<sub>m</sub></em>&nbsp;(proportion of misplaced edges, i.e. edges meant to be frustrated by construction). Moreover, we make the assumption that the proportion of misplaced edges is the same inside and between the modules. When it comes to incomplete unweighted signed networks, we introduce two more parameters, which are the density d of the graph and the proportion q neg of the negative edges. The last parameter q<sub>neg</sub>&nbsp;allows to control the ratio of positive to negative edges. For complete unweighted signed networks with d = 1, we generate 20 replications for parameter values l<sub>0</sub>&nbsp;= 3, n&nbsp;<em>&isin;</em>&nbsp;{32, 36, 40, 45, 50} and&nbsp;<em>q<sub>m</sub></em>&nbsp;<em>&isin;</em>&nbsp;{0.1, 0.2, 0.3, 0.4, 0.5, 0.6}. In these networks, the value of&nbsp;<em>q<sub>neg</sub>&nbsp;</em>with the considered parameters is approximately equal to 0.7. For incomplete unweighted signed networks with d&nbsp;<em>&isin;</em>&nbsp;{0.25, 0.50}, we generate 20 replications for parameter values l<sub>0</sub>&nbsp;= 3, n&nbsp;<em>&isin;</em>&nbsp;{32, 36, 40}, q<sub>m</sub>&nbsp;<em>&isin;</em>&nbsp;{0.1, 0.2, 0.3, 0.4, 0.5, 0.6} and&nbsp;<em>q<sub>neg</sub>&nbsp;</em><em>&isin;</em>&nbsp;{0.3, 0.5, 0.7}. In total, we produce 600 and 1,080 instances for complete and incomplete networks, respectively, which makes a total of 1,680 instances.</li> <li>partitions: folder containing the partitioning results of two methods: EnumCC(3) vs. OneTreeCC(). Note that the results of OneTreeCC() are not shown for space considerations, except for those with n=50.</li> <li>Results</li> </ul> <li>Experiments for Dataset 2</li> <ul> <li>benchmark netwoks: We generate these complete and incomplete networks through our random signed network generator, which is publicly available online. The optimal solution for a generated network is known by construction. For a given n, d and l<sub>0</sub>, we first create a perfectly structurally balanced (i.e., internally positive and externally negative) signed network with a built-in module structure. The underlying module structure constitutes the optimal partition. Then, in order to take into account different positive to negative ratio values for internal and external edges we generate several signed networks by perturbing the initial signed network without affecting its underlying optimal partition, thanks to its definition of stability range. We generate signed networks with parameter values n&nbsp;<em>&isin;</em>&nbsp;{30, 40, 50, 60, 70, 90}, d&nbsp;<em>&isin;</em>&nbsp;{0.25, 1.00} and l<sub>0</sub>&nbsp;<em>&isin;</em>&nbsp;{2, 4, 6}. In total, we produce 214 and 184 instances for complete and incomplete networks, respectively, which makes a total of 398 instances.</li> <li>benchmark partitions: folder containing the partitioning results of two methods: CoNS(<em>r<sub>max</sub></em>) with vs. without MVMO pruning, where&nbsp;<em>r<sub>max</sub></em>&nbsp;&isin; {3,4}.</li> <li>results: two `csv` files containing benchmark results between CoNS(<em>r<sub>max</sub></em>) with vs. without MVMO pruning, where r<sub>max</sub>&nbsp;&isin; {3,4}.</li> </ul> </ul> </div> </div> </div>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Data and Code for "Comparing the Effects of Euclidean Distance Matching and Dynamic Time Warping in the Clustering of COVID-19 Evolution"

<p>This repository contains the datasets and data sources, analysis code, and workflow associated with the manuscript "<em>Comparing the Effects of Euclidean Distance Matching and Dynamic Time Warping in the Clustering of COVID-19 Evolution</em>". The following resources are provided:</p> <ul> <li> <p><strong>Data Files</strong>:</p> <ul> <li><code>time_series_data.csv</code>: A curated time series dataset with dates as rows and NUTS 2 regions as columns. Each column is labeled using a 4-letter abbreviation format "CC.RR", where "CC" represents the country code and "RR" represents the region code. This same abbreviation is also included in the accompanying GeoJSON file.</li> <li><code>geometry_data.geojson</code>: A GeoJSON file representing the spatial boundaries of the NUTS 2 regions, with the same 4-letter abbreviations used in the CSV file. EPSG:4326.</li> <li><code>COVID19_data_sources.xlsx</code>: This Excel file contains important metadata regarding the sources of COVID-19 data used in this study. It includes: <ul> <li>Source of the data for each country</li> <li>Official website(s)</li> <li>The agency responsible for the data</li> <li>Description of the processing steps used to curate the data into the final time series.</li> </ul> </li> </ul> </li> <li> <p><strong>Code</strong>:</p> <ul> <li><code>analysis.py</code>: A Python script used to process and analyze the data. This code can be run using Python 3.x. The libraries required to run this script are listed in the first lines of the code. The code is organized in different numbered sections (1), (2), ... and sub-sections (1a), (1b) ... Make sure to run the script one (sub-)section at a time, so that everything stays overviewable and you don't get all the output at once.</li> </ul> </li> <li> <p><strong>Workflow</strong>:</p> <ul> <li><code>workflow.png</code> : A detailed workflow according to the Knowledge Discovery in Databases (KDD) process, outlining the steps involved in processing and analyzing the data, including the methods used. This workflow provides a comprehensive guide to reproducing the analysis presented in the paper.</li> </ul> </li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo36/100

"Filaments in and between galaxy clusters at low and mid-frequency with the SKA telescope" (Figures 6 to 13, B.1 and B.2)

<p>This file contains Figures 6 to 13, B.1 and B.2, from the paper "Filaments in and between galaxy clusters at low and mid-frequency with the SKA telescope", Vacca et al., accepted for publication on A&amp;A on 25 September 2024.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Genomic localization bias of secondary metabolite gene clusters and association with histone modifications in Aspergillus

<p>Table S4 (Distribution of Orthologous groups) associated with the publication 'Genomic localization bias of secondary metabolite gene clusters and association with histone modifications in Aspergillus' is deposited at Zenodo.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

JWST/NIRCam Images of the Orion Nebula Cluster

<p>Infrared mosaic images of the center of the Orion Nebula Cluster from NIRCam on JWST in 12 filters (F115W, F140M, F162M, F182M, F187N, F212N, F277W, F300M, F335M, F360M, F444W, F470N). The observations were performed by McCaughrean &amp; Pearson (2023) (arXiv.2310.03552). These images were reduced by Luhman (2024).</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Data for 'NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning'

<p>The uploaded files include two archives for <a href="https://pubs.acs.org/doi/10.1021/acs.jproteome.4c00300" target="_blank" rel="noopener">NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning</a>. The '<em>mgf_data</em>' archive contains all MGF files used in the study, while the '<em>sample_data</em>' archive includes sequencing data, clustering data generated using <code>MSCluster</code>, and XCorr calculation data computed with <code>CometX</code>, all of which were used in the research.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Clusters of topic modelling and naive Bayes classifier - FR CH newspapers

<p>Clusters of articles based on annotations produced by topic modelling and&nbsp;naive Bayes classifier applied to French language newspapers of Switzerland, published between 1900 and 1944 and containing the characters &quot;europ&quot;, extracted from the impresso app.&nbsp;</p>

opencc-by-4.0Jul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record