Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,782
datasets available to search
ShareScore release 0.7.1
Dataset results
1,782 results for “algorithms”
Data Sets for the study "What Performance Indicators to Use for Self-Adaptation in Multi-Objective Evolutionary Algorithms"
<p>This is the dataset of the paper "What Performance Indicators to Use for Self-Adaptation in Multi-Objective Evolutionary Algorithms"</p> <ul> <li>figures.ipynb provides postprocessing codes for using the attached data to generate tables and figures in the paper.</li> <li>the csv folder consists of the raw data of function evaluations that algorithms used to hit each solution the first time.</li> <li>the metric folder consists of the processed data, which records the convergence process of two single objectives (y1 and y2), Hypervolume, and the number of obtained Pareto solutions.</li> <li>the gif folder consists of the gif files plotting the convergence process of solving LOTZ.</li> </ul> <p> </p> <p>The code used to generate this data is available at https://github.com/FurongYe/GSEMO</p>
Parkinson disease validation algorithms
<p>The primary objective of this study was to validate two algorithms for the identification of persons with PD using clinical diagnosis as the reference standard on an Italian sample of people with PD. </p> <p>Two algorithms (index tests) applied to health administrative databases (hospital discharge, drug prescriptions, exemptions for medical costs) were validated against clinical diagnosis of PD by an expert neurologist (reference standard) in a cohort of consecutive outpatients.</p> <p>The two algorithms showed high accuracy for identifying patients with PD: one with greater sensitivity 94.2% (95% CI 88.4 – 97.6) and the other with greater specificity 98.1% (95% CI 97.7 – 98.5).</p>
AIRA (Atmospheric Rivers Identification Algorithm) input dataset and results
<p>This dataset contains the output of three regional climate simulations covering Europe for the period 1991-2010, used to apply Atmospheric Rivers Identification Algorithm (AIRA in Spanish), and the results of its performance. The simulations were carried out by the MAR group (www.um.es/gmar) of the University of Murcia, using the WRF-Chem model (v3.6.1). Each simulation includes different levels of aerosol interactions.</p>
Synthetic Datasets from the 2023 ECXAI Workshop Paper on Principled Benchmarking for Rule Set Learning Algorithms
<p>Synthetic datasets generated as part of the demonstration given in the paper <em>Towards Principled Synthetic Benchmarks for Explainable Rule Set Learning Algorithms</em> presented at the <em>Evolutionary Computing and Explainable Artificial Intelligence</em> (ECXAI) workshop taking place as part of the 2023 GECCO conference.</p>
Algorithm performance comparison in a classification task
<p>Results of algorithm performance comparison experiment.</p>
Assessment of the performance of the atmospheric correction algorithm MAJA for Sentinel-2 surface reflectance estimates
<p>Data associated to the paper "Assessment of the performance of the atmospheric correction algorithm MAJA for Sentinel-2 surface reflectance estimates", Colin, J. et al.</p> <p>Contact: jerome.colin[at]cnrs.fr<br> CESBIO Lab, Toulouse, France</p> <p>Content:<br> - APU_all_sites: APU plots for all the ACIX-II sites<br> - Maja_L2A_noadj_notopo: MAJA Level-2A subsets used to compare against ACIX-II reference reflectances for all sites and time steps<br> - quicklooks_all_sites: quicklooks for all sites</p> <p> </p>
Predicting species richness and diversity using satellite remote sensing and random forest machine learning algorithm
<p><strong>Aims</strong>: Remote sensing approaches could be beneficial for monitoring and compiling essential biodiversity data because it is cost-effective and allows for coverage of large areas over a short period. This study investigated the relationship between multispectral remote sensing data from Landsat 8 and Sentinel 2 and species richness and diversity in mountainous and protected grasslands.</p> <p><strong>Locations</strong>: Golden Gate Highlands National Park, Free State, South Africa. </p> <p><strong>Methods</strong>: In-situ data of plant species composition and cover from 142 plots with 16 releves each were distributed across the study site and used to calculate species richness and Shannon-wiener species diversity index (species diversity. We used a machine-learning random forest algorithm to optimise the prediction of species richness and diversity. The algorithm was used to identify the optimal spectral bands and vegetation indices for estimating species richness and diversity. Subsequently, the selected bands and vegetation indices were used to estimate species richness through random forest regression. </p> <p><strong>Results</strong>: This research found weak relationships between remote sensing vegetation indices and the diversity metrics, but significant relationships were found between some spectral bands and diversity metrics. Moreover, using machine learning random forest, the multispectral datasets exhibited strong predictive powers. In this investigation, for both sensors, near-infrared (NIR) seemed to be the most selected band to explain species diversity in mountainous grasslands.</p> <p><strong>Main</strong> <strong>conclusions</strong>: This finding further ascertains the efficiency of using NIR in vegetation mapping. This research shows that NIR, SAVI and EVI are the most adequate for predicting species richness and diversity in mountainous grasslands with relatively good accuracies.</p>
Input Data for: On the External Validity of Average-Case Analyses of Graph Algorithms
<p><strong>Data</strong></p> <p>All networks from <a href="https://networkrepository.com"><code>networkrepository.com</code></a> [1] with at most 1M edges (fall 2020). Weights and edge directions have been removed. For graphs with isomorphic largest connected component, only one copy has been kept.</p> <p>This is the raw data necessary to reproduce our experiments in <em>On the External Validity of Average-Case Analyses of Graph Algorithms</em>.</p> <p>[1] Ryan A. Rossi and Nesreen K. Ahmed, <em>The Network Data Repository with Interactive Graph Analytics and Visualization</em> (AAAI 2015)</p>
Dataset: Harmonic Integration via Randomized Kaczmarz Algorithms
<p>The dataset used for the analyses in our paper: "<a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4365845">On the Conjugate Symmetry and Sparsity of the Harmonic Decomposition of Parametric Surfaces with the Randomised Kaczmarz Method</a>".</p> <p>The codes are available on our <a href="https://github.com/eesd-epfl/harmonic-integration-via-kaczmarz">EESD GitHub webpage</a>.</p> <p>The attached files are all MAT files (Matlab 7.0). These files will be read and processed directly with our Python3.8 codes. The codes that generated these input data are referred to in our paper (via Spherical harmonics, spherical cap harmonics, and disk harmonics approaches).</p>
Data for training AMSR2-CNN and its corresponding machine learning algorithm
<p>Despite the availability of multiple decades of passive microwave measurements from satellite platforms, their utility for developing quantitative, spatially distributed estimates of snowpack is yet to be realized. A major bottleneck is the use of simple conceptual retrieval model formulations that are ineffective in representing the significant heterogeneity and complexity of snow evolution, particularly over areas with complex topography and forest regions. Here we demonstrate a physics-constrained and interpretable Convolutional Neural Network (CNN) to learn the functional relationship utilizing multi-channel passive microwave brightness temperature measurements from the Advanced Microwave Scanning Radiometer 2 (AMSR2) and in-situ snow depth observations. The machine learning approach with CNN generates vastly improved snow depth estimates relative to the standard AMSR2 estimates. Compared to independent in-situ measurements of snow depth over the Continental United States, the domain averaged Pearson correlation measure is three times higher than that of the standard AMSR2 estimates (R<sup>2</sup>: 0.68 versus 0.21), while the systematic errors are reduced by approximately fourfold. Further, the CNN-based snow depth estimates also exhibit notable enhancements in regions with forests, deep snow, and melting snow, thereby alleviating the limitations faced by traditional algorithms in retrieving accurate snow depths. The interpretation of the CNN framework further indicates that the machine learning approach dynamically leverages both volume scattering and emission components from a suite of measured passive microwave signals to generate more accurate snow depth retrievals. The results of this study provide an important benchmark of high-quality snow retrievals from passive microwave satellite measurements by maximizing their information content.</p>
A real dataset for evaluating hybrid clustering algorithm
<p>This is a real dataset for testing our hybrid clustering algorithm, and details can be found in our manuscript.</p>
Datasets S1 ~ S6 for evaluating hybrid clustering algorithm
<p>These datasets are used to evaluate our hybrid clustering algorithm. For details of our algorithm, please refer to https://github.com/junhaiqi/Hybrid_clustering.git.</p>
A Customized Bayesian Algorithm to Optimize Enzyme-Catalyzed Reactions
<p>Data underlying the figures in the publication "A Customized Bayesian Algorithm to Optimize Enzyme-Catalyzed Reactions" published in <em>ACS Sustain. Chem. Eng.</em>, <strong>2023</strong>, <a href="https://doi.org/10.1021/acssuschemeng.3c02402">https://doi.org/10.1021/acssuschemeng.3c02402</a>.</p> <p>Table of contents:</p> <ul> <li><strong>sc3c02402_si_001.pdf</strong>, <strong>sc3c02402_si_002.pdf</strong>, <strong>sc3c02402_si_003.pdf, sc3c02402_si_004.pdf</strong>: DNA sequences, supplementary figures, materials, methods, availability of the program, synthesis protocols</li> <li><strong>sc3c02402_raw_data.zip</strong>: Raw data</li> </ul>
Comparison of Maximal End Component Decomposition Algorithms: Data and Code
<p>Code and Data used for the Bachelor Thesis "Comparison of Maximal End Component Decomposition Algorithms"</p>
Experimental Results for the AAAI 2023 Paper "On Total-Order HTN Plan Verification with Method Preconditions -- An Extension of the CYK Parsing Algorithm"
<p>This is the collection of the experimental results reported in the paper "On Total-Order HTN Plan Verification with Method Preconditions -- An Extension of the CYK Parsing Algorithm". For a detailed explanation, please read the README.md file.</p>
Benchmarking splice variant prediction algorithms using massively parallel splicing assays
<p>Dataset, jupyter notebooks, and support python modules for "Benchmarking splice variant prediction algorithms using massively parallel splicing assays" (Smith and Kitzman, 2023)</p>
Modified LRphase and Simulated Dataset for "LRphase: an efficient algorithm for assigning haplotypic identity to long reads"
<p><strong>A modified version of LRphase was used to simulate reads for a single hypothetical human genome with maternal and paternal phasing information. Briefly, haplotype-specific reference sequences were generated with `bcftools consensus` <a href="https://paperpile.com/c/VO12CC/SIk9">(Li 2011)</a> based on the rescued GIAB VCF and hg38 human reference sequence. Each haplotype-specific fasta was fed separately into pbsim2 <a href="https://paperpile.com/c/VO12CC/NMmE">(Ono, Asai, and Hamada 2021)</a> as the reference from which simulated reads were randomly drawn, up to 1X coverage. Parameters controlling the read length distribution, sequencing, and base calling error rates were set to emulate typical performance of the MinIon sequencing platform with flow cell version R10.4.1 (https://nanoporetech.com/products/minion). These are as follows: `--depth 1 –hmm_model R103.model --difference-ratio '23:31:46' --length-mean 25000 --length-min 100 --length-max 1000000 –length-sd 20000 --accuracy-mean 0.98 --accuracy-min 0.01 --accuracy-max 1.00`. Simulated reads were aligned to the hg38 reference genome with minimap2 <a href="https://paperpile.com/c/VO12CC/Y2HP">(Li 2018)</a> and correct phasing and alignment coordinates were encoded in the read names. Finally, samtools <a href="https://paperpile.com/c/VO12CC/YLss">(Li et al. 2009)</a> was used to remove duplicated and supplementary reads, and concatenate, sort, and index reads into a single combined bam file. Of 258,539 total reads, 246,210 were mappable, and 178,504 overlapped at least one heterozygous variant in HG001.</strong></p>
Applicability of Search-based Algorithms for Model-based Game Play Testing - an Empirical Study
<p>Experimental data used for the study on EvoMBT approach for testing of games.</p>
Peak lists, assignments, and structures of 100 proteins determined by the ARTINA algorithm
<h2> </h2><p>The time-consuming and complex data analysis process constitutes a limitation for biomolecular NMR. It typically requires weeks or months of manual work of a trained expert to identify and interpret thousands of signals, recorded in a series of multidimensional NMR spectra, in order to determine the sequence-specific resonance assignments or the three-dimensional structure of a single protein. </p><p>To solve this problem, we developed ARTINA, a machine learning-based method that uses as input only NMR spectra and the protein sequence, and delivers signal positions, resonance assignments, and structures strictly without any human intervention. Tested on a 100-protein benchmark comprising 1329 multidimensional NMR spectra, ARTINA demonstrated its ability to solve structures with 1.44 Å median RMSD to the PDB reference and to identify 91.36% correct NMR resonance assignments. </p><p>ARTINA is freely accessible for non-commercial users at NMRtist (https://nmrtist.org), an online platform that combines deep learning, large-scale optimization, and cloud computing to offer full automation of NMR spectra analysis. Our website provides virtual storage for NMR spectra deposition together with a set of applications designed for automated peak picking, chemical shift assignment, and protein structure determination. The system can be used by non-experts and without IT infrastructure on the user's side, allowing protein NMR spectra interpretation within hours after completion of the measurements. With NMRtist, the effort for a protein assignment or structure determination by NMR essentially reduced to the preparation of the sample and the spectrum measurements. </p>
A Development of Inflammatory Bowel Disease Pattern Identification Algorithm Using Case Series Data
ClinicalTrials.gov study NCT04296500. IPD Sharing: Not stated. Countries: 1. Publications: 6.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.