Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,782
datasets available to search
ShareScore release 0.7.1
Dataset results
1,782 results for “algorithms”
Different orthology inference algorithms generate similar predicted orthogroups among Brassicaceae species
Open the record for dataset details and reuse information.
Data from: Illusory speeding-up and slowing-down of objects moving at constant speed emerges from natural motion detection algorithms
Open the record for dataset details and reuse information.
Path-finding algorithm as a dispersal assessment method for invasive species with human-vectored long-distance dispersal event
Open the record for dataset details and reuse information.
Tassie BRUV: A benchmark data set for computer vision and movement quantification algorithms
Open the record for dataset details and reuse information.
Cloud_ICA: A deterministic cloud-overlap algorithm for generating a complete set of independent column atmospheres
Open the record for dataset details and reuse information.
Assessing predictive performance of supervised machine learning algorithms for a diamond pricing model
Open the record for dataset details and reuse information.
NH and ME Landsat chlorophyll-a retrieval algorithms and in situ measurements 2000 (Landsat 7), 2013-2015 (Landsat 8)
Predicting algal blooms has become a priority for scientists, municipalities, businesses, and citizens. Remote sensing offers solutions to the spatial and temporal challenges facing existing lake research and monitoring programs that rely primarily on high-investment, in situ measurements. Techniques to remotely measure chlorophyll-a (chl-a) as a proxy for algal biomass have been limited to specific large water bodies in particular seasons and narrow chl-a ranges. Thus, a first step toward prediction of algal blooms is generating regionally robust algorithms using in situ and remote sensing data. This study explores the relationship between in-lake measured chl-a data from Maine and New Hampshire lakes and remotely-sensed chl-a retrieval algorithm outputs. Landsat 8 images were obtained and then processed after required atmospheric and radiometric corrections. Six previously developed algorithms were tested on a regional scale on eleven scenes from 2013-2015 covering 192 lakes. Additionally, data related to one Landsat 7 scene (2000) are included in this data set. Boucher, J, K.C. Weathers, H. Norouzi, B. Steele. In Press. Assessing the effectiveness of Landsat 8 chlorophyll-a retrieval algorithms for regional freshwater monitoring. Ecological Applications.
Control algorithm simulation results
<p>Torque vs. Speed maps of the 200kW pure SynRel power settings </p>
Project O2 - A Cooperative Bypassing Algorithm for Connected and Autonomous Vehicles in Mixed Traffic
<p>The dataset includes Python code and VISSIM file for the research paper "A Cooperative Bypassing Algorithm for Connected and Autonomous Vehicles in Mixed Traffic".</p> <p> </p> <p> </p>
Experimental Data Sets for the study "Benchmarking a $(\mu+\lambda)$ Genetic Algorithm with Configurable Crossover Probability"
<p>This is the experimental result of the study "Benchmarking a (μ+λ) Genetic Algorithm with Configurable Crossover Probability". A novel (μ+λ) GA is proposed and benchmarked, in which we stochastically determine whether to apply the crossover operator either for each individual or generation with a crossover probability <span class="math-tex">\(p_c\)</span>. This data set consists of two parts:</p> <ol> <li>The results of (μ+λ) GA on 25 pseudo-Boolean problems defined in <em>IOHprofiler </em>(<a href="https://iohprofiler.github.io/">https://iohprofiler.github.io/</a>) with the following setup: <span class="math-tex">\(\mu \in \{10, 50, 100\}, \lambda \in \{1, \lceil\mu/2\rceil, \mu\}, p_c\in\{0, 0.5\}.\)</span> <ul> <li>'IOHprofiler_Problems_standard_bit_mutation.csv' --> the (μ+λ) GA with standard bit mutation.</li> <li>'IOHprofiler_Problems_fast_mutation.csv' --> the (μ+λ) GA with fast mutation.</li> </ul> </li> <li>The results of (μ+λ) GA on OneMax and LeadingOnes problems with the following setup: <span class="math-tex">\(n \in \{64,100,150,200,250,500\}, \mu \in \{2,3,5,8,10,20,30,...,100\}, \\ \lambda \in \{1, \lceil \mu/2 \rceil, \mu\}, \text{and }p_c \in \{0.1 k \mid k \in [0..9]\}\cup\{0.95\}.\)</span> <ul> <li>'OneMax_raw.csv' --> the fixed-target running time/first hitting time from 100 independent runs for target values in <span class="math-tex">\([1..n]\)</span>.</li> <li>'OneMax_summary.csv' --> the mean, median, standard deviation, some quantiles, expected running time (ERT), the number of successful runs, and the success rate from 100 independent runs for target values in <span class="math-tex">\([1..n]\)</span>.</li> <li>'LeadingOnes_raw.csv' --> the same with 'OneMax_raw.csv' for LeadingOnes.</li> <li>'LeadingOnes_summary.csv' --> the same with 'OneMax_summary.csv' for LeadingOnes.</li> </ul> </li> </ol> <p><strong>Contact</strong>: if you have any questions or suggestions, please feel free to contact <a href="https://www.universiteitleiden.nl/en/staffmembers/furong-ye#tab-1">Furong Ye</a> or <a href="http://www-ia.lip6.fr/~doerr/">Carola Doerr</a>.</p>
The Locus Algorithm Local Catalogue
<p>This is the Local Catalogue generated as an intermediate step in the use of the Locus Algorithm to produce catalogues of optimized pointings for differential photometry of various classes of object. This catalogue mirrors the format of the SDSS Object catalogue but contains significantly reduced data to reduce the I/O demands on the system.</p>
An Experimental Study of Operator Choices in the (1+(λ,λ)) Genetic Algorithm
<p>This dataset contains experimental results that accompanies the same-named paper accepted to the MOTOR conference and scheduled to be published in a CCIS volume.</p> <p>The following data files are a part of this dataset, whose meaning should become clear after reading the paper:</p> <ul> <li>w1-runs.csv: the results of each of the runs, each of the configurations on the OneMax problem;</li> <li>w2-runs.csv: same for the LinInt<sub>2</sub> problem;</li> <li>w5-runs.csv: same for the LinInt<sub>5</sub> problem;</li> <li>ms-runs.csv: same for the easy random MAX-SAT instances;</li> <li>w1-stats.csv: the quartile statistics for the OneMax problem;</li> <li>w2-stats.csv: same for the LinInt<sub>2</sub> problem;</li> <li>w5-stats.csv: same for the LinInt<sub>5</sub> problem;</li> <li>ms-stats.csv: same for the easy random MAX-SAT instances;</li> <li>tunings-static.csv: the outcomes of irace for the static configurations (with fixed λ);</li> <li>tunings-dynamic.csv: the outcomes of irace for the dynamic configurations (with self-adjusting λ).</li> </ul> <p>Apart from that, the file pictures.pdf presents the contents of *-stats.csv files visually in a concide way.</p>
Genetic algorithm-based personalized models of human cardiac action potential
<p>We present a novel modification of genetic algorithm (GA) which determines personalized parameters of cardiomyocyte electrophysiology model based on set of experimental human action potential (AP) recorded at different heart rates. In order to find the steady state solution, the optimized algorithm performs simultaneous search in the parametric and slow variables spaces. We demonstrate that several GA modifications are required for effective convergence. Firstly, we used a mutation operator, based on Cauchy amplitude distribution along with a random direction in the parametric space. Secondly, relatively large number of elite organisms (6-10 % of the population passed on to new generation) was required for effective convergence. Test runs with synthetic AP as input data indicate that algorithm error is low for high amplitude ionic currents (1.6±1.6% for IKr, 3.2±3.5% for IK1, 3.9±3.5% for INa, 8.2±6.3% for ICaL). Experimental signal-to-noise ratio above 28 dB was required for high quality GA performance. GA was validated against optical mapping recordings of human ventricular AP and mRNA expression profile of donor hearts. In particular, GA output parameters were rescaled proportionally to mRNA levels ratio between patients. We have demonstrated that mRNA-based models predict the AP waveform dependence on heart rate with high precision. The latter also provides a novel technique of model personalization that makes it possible to map gene expression profile to cardiac function. </p>
Seafloor Density Measurements, Prediction, and Associated Uncertainty for "Predicting global marine sediment density using the random forest regressor machine learning algorithm"
<p>Global seafloor density prediction results using the random forest regressor machine learning algorithm. </p> <p>Dataset S1. Seafloor density measurements. Columns are labeled with a header and include associated drilling project and measurement type for each sample. File format: CSV text file</p> <p>Dataset S2. Seafloor density prediction results from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p> <p>Dataset S3. Seafloor density prediction standard deviation from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p>
Data for the paper "An Inversion Algorithm for Deriving Shallow Structure using Co-located Wind, Pressure, and Seismic Data"
<p>Wind, pressure, and seismic data used in the paper "An Inversion Algorithm for Deriving Shallow Structure using Co-located Wind, Pressure, and Seismic Data"</p>
NEMO, HIDRA and Tide Gauge Datasets for HIDRA Machine Learning Algorithm Verification
<p>Supporting sea level datasets for paper:</p> <p>"HIDRA 1.0: Deep-Learning-Based Ensemble Sea Level Forecastingin the Northern Adriatic"</p> <p>by Lojze Žust, Anja Fettich, Matej Kristan, and Matjaž Ličer</p>
Standard Bouguer anomaly model achieved by multi-source Bouguer gravity anomaly Bayesian data fusion algorithm in Sichuan-Yunnan region
<p>* Method: Based on the equivalent source inversion and Bayesian uncertainty quantization theory, a new multi-source gravity data fusion algorithm is developed, which effectively solves the multi-source data fusion problem with different noise and datum.</p> <p>* Standard Bouguer anomaly is Fused from WGM2012 Bouguer gravity anomaly model and 394 gravity profile data measured in Sichuan-Yunnan region. Fusion anomaly results can eliminate datum draft between multi-source gravity and reduce incoherent noise.</p> <p>* Spatial resolution of the standard Bouguer anomaly is about 20 kilometers.</p> <p>* Correcting deviations means the difference between the fused standard Bouguer anomaly model and the WGM2012 Earth gravity model.</p>
The Camouflage Machine: Optimising protective colouration using deep learning with genetic algorithms
Evolutionary biologists frequently wish to measure the fitness of alternative phenotypes using behavioural experiments. However, many phenotypes are complex. For example colouration: camouflage aims to make detection harder, while conspicuous signals (e.g. for warning or mate attraction) require the opposite. Identifying the hardest and easiest to find patterns is essential for understanding the evolutionary forces that shape protective colouration, but the parameter space of potential patterns (coloured visual textures) is vast, limiting previous empirical studies to a narrow range of phenotypes. Here we demonstrate how deep learning combined with genetic algorithms can be used to augment behavioural experiments, identifying both the best camouflage and the most conspicuous signal(s) from an arbitrarily vast array of patterns. To show the generality of our approach, we do so for both trichromatic (e.g. human) and dichromat (e.g. typical mammalian) visual systems, in two different habitats. The patterns identified were validated using human participants; those identified as the best for camouflage were significantly harder to find than a tried-and-tested military design, while those identified as most conspicuous were significantly easier than other patterns. More generally, our method, dubbed the 'Camouflage Machine', will be a useful tool for identifying the optimal phenotype in high dimensional state-spaces.
Identifying Energy Efficiency Patterns in Sorting Algorithms via Abstract Syntax Tree Mining
<p>Replication package for Submission "Identifying Energy Efficiency Patterns in Sorting Algorithms via Abstract Syntax Tree Mining".</p> <p>Authors kept anonymous for review.</p>
Synthetic realistic noise-corrupted PPG database and noise generator for the evaluation of PPG denoising and delineation algorithms
<p><strong>Overview </strong></p> <p>This database is meant to evaluate the performance of denoising and delineation algorithms for PPG signals affected by noise. The noise generator allows applying the algorithms under test to an artificially corrupted reference PPG signal and comparing its output to the output obtained with the original signal. Moreover, the noise generator can produce artifacts of variable intensities, permitting the evaluation of the algorithms' performance against different noise levels. The reference signal is a PPG sample of a healthy subject at rest during a relaxing session.</p> <p> </p> <p><strong>Database</strong></p> <p>The database includes 1 recording of 72 seconds of synchronous PPG and ECG signals sampled at 250 Hz using a Medicom device, ABP-10 module (Medicom MTD Ltd., Russia). It was collected from a healthy subject during an induced relaxation by guided autogenic relaxation. For more information about the data collection, please refer to the following publication: <a href="https://pubmed.ncbi.nlm.nih.gov/30094756/">https://pubmed.ncbi.nlm.nih.gov/30094756/</a></p> <p>In addition, PPG signals corrupted by the noise generator at different levels are also included in the database.</p> <p> </p> <p><strong>Realistic noise generator</strong></p> <p>Motion Artifacts in PPG signals generally appear in the form of sudden spikes (in correspondence to the subject's movement) and slowly varying offsets (baseline wander) due to the changes in distance between the skin and the sensor after every sudden movement. For this reason, conventional noise generators — using random noise drawn from different distributions such as Gaussian or Poissonian — do not allow to properly evaluate the algorithm's performance, as they can only provide unrealistic noises compared to the one commonly found in PPG signals. To overcome this issue, we designed a more realistic synthetic noise generator that can simulate those two behaviors, enabling us to corrupt a reference signal with different noise levels. The details about noise generation are available in the reference paper.</p> <p> </p> <p><strong>Data Files</strong></p> <p>The reference PPG signal can be found in <em>Datasets\GoodSignals\PPG</em> and the simultaneously acquired ECG in <em>Datasets\GoodSignals\ECG</em>. The folder <em>Datasets\NoisySignals</em> contains 340 noisy PPG signals affected by different levels of noise. The names describe the intensity of the noise (evaluated in terms of the standard deviation of the random noise used as input for the noise generator, see reference paper). Five noisy signals are produced for every noise level by running the noise generator with five random seeds each (for noise generation).</p> <p>Name convention: <em>ppg_stdx_y</em> denotes the y-th noisy PPG signal produced using a noise with a standard deviation of x.</p> <p><em>Datasets\BPMs</em> contains the ground truth for the heart-rate estimation computed in windows of 8s with an overlap of 2s.</p> <p><strong>Code</strong></p> <p>The folder <em>Code </em>contains the MATLAB scripts to generate the noisy files by generating the realistic noise with the function noiseGenerator.</p> <p><strong>When referencing this material, please cite:</strong></p> <p>Masinelli, G.; Dell'Agnola, F.; Valdés, A.A.; Atienza, D. SPARE: A Spectral Peak Recovery Algorithm for PPG Signals Pulsewave Reconstruction in Multimodal Wearable Devices. <em>Sensors</em> <strong>2021</strong>, <em>21</em>, 2725. <a href="https://doi.org/10.3390/s21082725">https://doi.org/10.3390/s21082725</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.