Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
155
datasets available to search
ShareScore release 0.9.0
Dataset results
155 results for “Cluster analysis”
Dataset of paper "Applying Density-Based Clustering for the Analysis of Emission Events in Real Driving Emissions"
<p>This dataset includes the signal traces for the events used in the publication "Applying Density-Based Clustering for the Analysis of Emission Events in Real Driving Emissions Calibration", MDPI Future Transportation, 2024 (doi: 10.3390/futuretransp4010004).</p> <p>The data is available in a frequency of 1 Hz and the signals are arranged in separate tables. Please note, that the order of the events is not consistent for the individual signals.</p> <p>The general vehicle specifications are listed below. For further information please refer to the paper.</p> <table> <tbody> <tr> <td><strong>Characteristic</strong></td> <td><strong>Unit</strong></td> <td><strong>Value</strong></td> </tr> <tr> <td>Vehicle weight</td> <td>kg</td> <td>> 2000</td> </tr> <tr> <td>Fuel</td> <td>-</td> <td>Gasoline</td> </tr> <tr> <td>Engine type</td> <td>-</td> <td>Turbo-charged 8 cylinder</td> </tr> <tr> <td>Engine power and torque</td> <td>kW / Nm</td> <td>> 400 / > 600</td> </tr> <tr> <td>Cubic capacity</td> <td>cm^3</td> <td>~ 4000</td> </tr> <tr> <td>Transmission</td> <td>-</td> <td>Automatic transmission (AT)</td> </tr> <tr> <td>Drivetrain</td> <td>-</td> <td>All-wheel drive (AWD)</td> </tr> <tr> <td>Exhaust aftertreatment system (EATS)</td> <td>-</td> <td>Three-way catalytic converter (TWC) and gasoline particulate filter (GPF)</td> </tr> <tr> <td>Condition of EATS</td> <td>-</td> <td>Stabilized EATS (~ 70 % of tests) and aged EATS (~ 30 % tests)</td> </tr> <tr> <td>Emission target</td> <td>-</td> <td>EU6d</td> </tr> </tbody> </table> <p> </p>
Data and analysis for: "Persistent Spatial Clustering and Predictors of Pediatric La Crosse Virus Neuroinvasive Disease Risk in Eastern Tennessee and Western North Carolina, 2003–2020"
<p>This is the initial release of the data and code corresponding to the manuscript submitted to PLoS Neglected Tropical Diseases. <strong>Please refer to the README.md file</strong> for a description of the contents of this repository and how to use them. The README file can be opened with a text editor, or viewed directly in the GitHub repository. The data and code are provided within a project directory with a reproducible R package library for ease and accuracy of reproducibility. </p> <p><strong>Ethics Approval</strong></p> <p>This study was approved by the University of Tennessee, Knoxville Institutional Review Board (UTK IRB-22-07079-XP) and the Tennessee Department of Health Institutional Review Board (TDH IRB 2021-0314). Data provided here is de-identified and aggregated (both temporally and spatially) to protect the privacy of individuals included in the study, in concordance with IRB and Data Use Agreements.</p>
ACCESS-AM2 Southern Ocean cloud and radiation data for k-means clustering and analysis
<p>The ACCESS-AM2 (Australian Community Climate and Earth-System Simulator - Atmospheric Model Version 2) data and k-means analysis used for the study described in Fiddes et al. 2022 '<em>Southern Ocean cloud and shortwave radiation biases in a nudged climate model simulation: does the model ever get it right?' .</em> </p> <p>Included files: </p> <ul> <li>modis_cluster_centres_2015-2019.nc - kmeans derived cluster centres for MODIS</li> <li>modis_cluster_labels_2015-2019.nc - kmeans derived cluster labels for MODIS </li> <li>bx400_cluster_labels_2015-2019.nc - kmeans fitted cluster label for model </li> <li>COSP_vars_bx400_2015-2019.nc - model data for analysis </li> </ul> <p>The code that performs the analysis/generates this data and has instructions for where to download MODIS data can be found here: https://github.com/sfiddes/code_for_publications_2022/tree/main/ACCESS_cloud_radiation_eval</p>
Dataset and models of TMLR 2024 Paper "Identifying and Clustering Counter Relationships of Team Compositions in PvP Games for Efficient Balance Analysis"
<p>This is a part of dataset and models of the paper published in TMLR 2024 (Transactions on Machine Learning Research, <a href="https://jmlr.org/tmlr/" target="_blank" rel="noopener">https://jmlr.org/tmlr/</a>).</p> <p>Including training datasets, testing datasets, and models.</p> <p>The example program for using this file will be put on the author's github repo branch: <a href="https://github.com/DSobscure/cgi_drl_platform/tree/game_balance_measures_tmlr" target="_blank" rel="noopener">https://github.com/DSobscure/cgi_drl_platform/tree/game_balance_measures_tmlr</a></p> <p> </p>
Supplementary Material for "A comparative high-resolution spectroscopic analysis of in situ and accreted globular clusters"
<p>This is a file containing supplementary material for the paper <em>A comparative high-resolution spectroscopic analysis of in situ and accreted globular clusters.</em> For each star in target globular clusters, it lists crucial information on the linelist analyzed. In particular:</p> <ol> <li>Star ID.</li> <li>Chemical element.</li> <li>Wavelength.</li> <li>log <em>gf</em></li> <li>Excitation potential.</li> <li>Measured equivalent width with uncertaintiy.</li> </ol>
Artifact Description/Artifact Evaluation/Computational Artifact for paper, entitled Analytic Roofline Modeling and Energy Analysis of the LULESH Proxy Application on Multi-Core Clusters
We provide reproducibility initiative dependencies (Artifact Description or Artifact Evaluation or Computational Results Analysis) appendix. To allow a third party to duplicate the findings, this article provides our extensive performance data artifact and describes further details regarding the software environments, experimental design, and methodology employed for the results shown in the paper. The computational artifacts will enable experienced performance engineers to reproduce and interpret the data shown in the paper in the appropriate way and to follow the conclusions we draw from it.
Bioenergetic cluster analysis – mitochondrial respiratory control in human fibroblasts
<p>Gnaiger E (2021) Bioenergetic cluster analysis – mitochondrial respiratory control in human fibroblasts. MitoFit Preprints 2021.8. doi:10.26124/mitofit:2021-0008 - https://www.mitofit.org/index.php/Gnaiger_2021_MitoFit_BCA</p> <p>All respirometric data that were used for meta-analysis are obtained from the original publications, were converted to SI units, and are available here as a basis for bioenergetic cluster analysis. Inverted regression analysis is illustrated by an example (Figures 1c and d).</p>
Multivariate analysis of FcR-mediated NK cell functions identifies unique clustering among humans and rhesus macaques - dataset
<p>Dataset from Tuyishime M, Spreng RL, et al. Multivariate analysis of FcR-mediated NK cell functions identifies unique clustering among humans and rhesus macaques. Frontiers in Immunology 2023 doi: 10.3389/fimmu.2023.1260377</p>
Characteristic and spatiotemporal variation of air pollution in Northern China based on correlation analysis and clustering analysis of five air pollutants
<p>original daily data for 'Characteristic and spatiotemporal variation of air pollution in Northern China based on correlation analysis and clustering analysis of five air pollutants'</p>
List of VLFEs obtained in the paper "Influence of a subducted oceanic ridge on the distribution of shallow VLFEs in the Nankai Trough as revealed by moment tensor inversion and cluster analysis"
<p>List of VLFEs obtained in Toh et al., (2020, GRL).</p> <p>"Influence of a subducted oceanic ridge on the distribution of shallow VLFEs in the Nankai Trough as revealed by moment tensor inversion and cluster analysis" by Akiko Toh, Wan-Jou Chen, Nozomu Takeuchi, Douglas Dreger, Wu-Cheng Chi, and Satoshi Ide. </p> <p> </p>
Dataset of structural and energetic descriptors for optimized geometries of the pyrene dimer in the S1 state and Jupyter Notebbok used for the unsupervised clustering, analysis and visualization
<p>Geometrical and energy data extracted from a set of 188 optimized geometries for the pyrene dimer in the first excited singlet state at the TD-CAMB3LYP + D3BJ / 6-31G* / C-CPCM(Cyclohexane) level. </p> <p>Coordinates (in xyz format) of the 188 optimized geometries considered.</p> <p>Jupyter notebook used to perform unsupervised clustering and analysis of the available structures.</p>
Supporting dataset for: "Plasma essential amino acid concentration and profile are associated with performance of lactating dairy cows as revealed through meta-analysis and hierarchical clustering"
<p>This dataset was used in the meta-analysis and hierarchical clustering published in "Plasma essential amino acid concentration and profile are associated with performance of lactating dairy cows as revealed through meta-analysis and hierarchical clustering" in the Journal of Dairy Science. We searched Web of Science and Google Scholar databases through March 2020 with the terms “plasma EAA,” “milk urea” or “blood urea,” and “dairy” or lactating dairy”. To be included in our study, the papers must have met the following selection criteria: (1) been published in English in a peer-reviewed journal; (2) reported dietary ingredients on a DM basis and at minimum dietary CP concentration; (3) used treatments based on diet changes (e.g., no infusion trials were included); (4) reported DMI, lactation performance, and milk components yield; (5) reported all individual [EAA]p (excluding Trp); and (6) reported blood urea-N or plasma urea-N. Infusion studies were excluded to avoid possible effects of method of EAA supply (e.g., infusion vs. feeding) and to narrow the scope of application. The final dataset included 22 studies and 96 dietary treatments. For a more complete description of the methods, please refer to the published paper. </p>
Characterizing Measures for the Assessment of Cluster Analysis and Community Detection
<p><strong>Description. </strong>The dataset is constituted of:</p> <ul> <li>`figs.zip`: an archive containing the plot files;</li> <li>`data&results.zip`: an archive containing the necessary data to perform our analysis, as well as result files.</li> </ul> <p>These are the resources used in the following articles:</p> <ol> <li>N. Arınık, V. Labatut and R. Figueiredo, "Characterizing measures for the assessment of cluster analysis and community detection", Modèles & Analyse des Réseaux : Approches Mathématiques & Informatiques (MARAMI), 2020. ⟨<a href="https://hal.archives-ouvertes.fr/hal-02993542">hal-02993542</a>⟩</li> <li>N. Arınık, R. Figueiredo, and V. Labatut, “Characterizing and comparing external measures for the assessment of cluster analysis and community detection,” <em>IEEE Access </em>9:20255–20276, 2021. DOI: <a href="http://doi.org/10.1109/access.2021.3054621">10.1109/access.2021.3054621</a> ⟨<a href="https://hal.archives-ouvertes.fr/hal-03124118">hal-03124118</a>⟩</li> </ol> <p><strong>Source code. </strong>The associated source code is available on GitHub: <a href="https://github.com/CompNet/ExtMeasEval">https://github.com/CompNet/ExtMeasEval</a></p> <p><strong>Citation. </strong>If you use these data, please cite the paper [2].</p> <p><br><code>@Article{Arinik2021,</code><br><code> author = {Arınık, Nejat and Figueiredo, Rosa and Labatut, Vincent},</code><br><code> title = {Characterizing and Comparing External Measures for the Assessment of Cluster Analysis and Community Detection},</code><br><code> journal = {IEEE Access},</code><br><code> year = {2021},</code><br><code> volume = {9},</code><br><code> pages = {20255-20276},</code><br><code> doi = {10.1109/access.2021.3054621},</code><br><code>}</code></p>
Spectral Cluster Supertree: Analysis Data
<p>Contains all datasets used in the Spectral Cluster Supertree paper. The datasets are composed of a set of rooted model trees, and rooted source trees to predict them. Please cite the appropriate papers, depending on which of the datasets you use.</p> <p>The <code>birth_death</code> folder contains our own dataset generated for our paper (where the generation process is explained), it aims to mimic what may be seen through divide and conquer algorithms for phylogenetic reconstruction. Parameters used to simulate an alignment were simulated under parameters estimated from a sequence alignment of 3 bacterial species (Kaehler et al., 2015) - see <code>alignment</code> folder.</p> <p>The <code>SMIDGenOutgrouped</code> folder contains both the SMIDGenOG (Fleischauer and Böcker, 2016) and SMIDGenOG-5500 dataset (Fleischauer and Böcker, 2017).</p> <p>The <code>SuperTriplets</code> folder contains the SuperTriplets dataset (Ranwez et al, 2010).</p> <p> </p>
Fig. 5. A in Morphometric Analysis And Interrelationship Of Seven Indonesian Hornbill Species (Aves, Bucerotidae) Utilizing Principal Component And Cluster Analysis
Fig. 5. A dendrogram illustrating the relationships among the seven Indonesian hornbill species based on 14 morphometric characters, constructed using the Average Linkage model. Legend: Aa = Anthracoceros albirostris, Am = Anthracoceros malayanus, Ru = Rhyticeros undulatus, Rp = Rhyticeros plicatus, Ac = Aceros cassidix, Br = Buceros rhinoceros, dan Bb = Buceros bicornis.
Fig. 1 in Morphometric Analysis And Interrelationship Of Seven Indonesian Hornbill Species (Aves, Bucerotidae) Utilizing Principal Component And Cluster Analysis
Fig. 1. Hornbill genus grouping based on a combination of body length characters (PC1) and beak characters (PC3): A — genus Rhyticeros; B — genus Buceros; C — genus Anthracoceros.
Fig. 3 in Morphometric Analysis And Interrelationship Of Seven Indonesian Hornbill Species (Aves, Bucerotidae) Utilizing Principal Component And Cluster Analysis
Fig. 3. The combination of tail length and head length of two hornbill species within the genus Anthracoceros.
Fig. 4 in Morphometric Analysis And Interrelationship Of Seven Indonesian Hornbill Species (Aves, Bucerotidae) Utilizing Principal Component And Cluster Analysis
Fig. 4. The combination of head length and tail length of three hornbill species within the genus Rhyticeros.
Fig. 3 in Cluster Analysis of Non-conserved Proteins of Trypanosoma cruzi Reference Strains Displays Parity between these Groupings (Peptidemes) and the Consensually Accepted Parasite Lineages
Fig. 3. Phenogram of the peptidemes (P) of eight Trypanosoma cruzi reference strains obtained using the SM coefficient and the UPGMA clustering algorithm, based on data from non-conserved proteins, as seen in SDS-PAGE analysis. The major peptidemes are indicated as mP 1 and mP 2. Their subgroups are identified on the right (P II, P VI, P I), and were numbered following their respective genetic types (TcII, TcVI, TcI), as currently used.
Fig. 1 in Cluster Analysis of Non-conserved Proteins of Trypanosoma cruzi Reference Strains Displays Parity between these Groupings (Peptidemes) and the Consensually Accepted Parasite Lineages
Fig. 1. Total protein profiles of eight Trypanosoma cruzi reference strains separated in 10% SDS-PAGE at 250 V, 25 mA, 90 min, and stained by Coomassie brilliant blue. The position of some conserved proteins is indicated on the right. M: molecular mass markers. (kDa) are indicated on the left.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.