Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
369
datasets available to search
ShareScore release 0.9.0
Dataset results
369 results for “Datasets, Benchmarking”
Performance Measurement Datasets of the HPC Benchmarks LAMMPS, MiniFE, LULESH for Hardware Counter Variance Analysis
Open the record for dataset details and reuse information.
Targeting protein-ligand neosurfaces using a generalizable deep learning approach [benchmark dataset]
<p>PDB files and processed surface meshes for the binder recovery benchmark. For the larger PDBbind decoy set only the PDB files are provided.</p>
Bins with greater than 50% completeness and less than 10% contamination from the paper "Benchmarking Metagenomic Binning Tools on Real Datasets Across Sequencing Platforms and Binning Modes".
<p>These bins, which exhibit greater than 50% completeness and less than 10% contamination, including MQ, NC, and HQ bins, were recovered from seven different data-binning combinations across five real-world datasets.</p> <p>CheckM2 results for the bins generated from each sample were also provided to corroborate the findings.</p> <p>Completeness and contamination were evaluated using CheckM 2 (version 1.0.2). We noted that the results from running CheckM2 may exhibit slight variations, but these differences do not affect the overall assessment.</p>
Model Comparison Benchmark Datasets, Results, and Peptide MD Trajectory
<p>Contents of this database include:</p> <ul> <li>aib9.out.zip: zipped file containing raw output of the Aib9 OpenMM MD simulation</li> <li>aib9_openmm.zip: zipped folder containing all input files for the Aib9 OpenMM MD simulation, as well as the generated trajectory</li> <li>data_input.zip: zipped folder containing input datasets for the GMM and Aib9 numerical experiments</li> <li>data_output.zip: zipped folder containing output data and experimental results for three generative models (NS, CFM, DDPM) on the input datasets, included for posterity (file names do not synchronize with most recent repository naming conventions)</li> </ul> <p>The files are automatically retrieved and unzipped via the data_accessor notebook in the GitHub repository.</p> <p> </p>
11 Benchmark Clean-Clean ER datasets in CSV format
<p>Contains:</p> <ul> <li><strong>D1</strong>: Contains restaurant descriptions, first introduced in OAEI 2010.</li> <li><strong>D2</strong>: Includes duplicate products from Abt.com and Buy.com.</li> <li><strong>D3</strong>: Matches product descriptions from Amazon and Google Base.</li> <li><strong>D4</strong>: Compares bibliographic data from DBLP and ACM.</li> <li><strong>D5, D6, D7</strong>: Contain descriptions of television shows and movies from TheTVDB, IMDb, and TMDb.</li> <li><strong>D8</strong>: Matches product descriptions from Walmart and Amazon.</li> <li><strong>D9</strong>: Involves bibliographic data from DBLP and Google Scholar.</li> <li><strong>D10</strong>: Links movie descriptions from IMDb and DBpedia.</li> <li><strong>D11</strong>: A large-scale dataset with millions of heterogeneous entities from two DBpedia versions spanning a 3-year gap.</li> </ul>
Benchmark Datasets CellCnn
<p>Benchmark datasets referenced in https://github.com/eiriniar/CellCnn</p>
AMR Benchmarking dataset - Assemblies
<p>Benchmarking dataset for AMR detection pipelines for assemblies focusing on ESKAPE pathogens in addition to Salmonella. This dataset consists of closed genomes from NCBI where paired-end Illumina data was available. The closed genomes from NCBI are provided in addition to assemblies based on the raw uploaded to NCBI, along with assemblies based on just the reads which mapped correctly to the closed assemblies. Reads were assembled using shovill v. 1.1.0 using SKESA and SPADES. Variants were called using snippy v. 4.6.0 and the mapped reads were extracted bedtools bamtofastq v2.29.2.</p>
AMR Benchmarking dataset - Mapped ReadSets - 1
<p>Benchmarking dataset for AMR detection pipelines for assemblies focusing on ESKAPE pathogens in addition to Salmonella. This dataset consists of closed genomes from NCBI where paired-end Illumina data was available. The closed genomes from NCBI are provided in addition to assemblies based on the raw uploaded to NCBI, along with assemblies based on just the reads which mapped correctly to the closed assemblies. Reads were assembled using shovill v. 1.1.0 using SKESA and SPADES. Variants were called using snippy v. 4.6.0 and the mapped reads were extracted bedtools bamtofastq v2.29.2.</p>
AMR Benchmarking dataset - Mapped ReadSets - 4
<p>Benchmarking dataset for AMR detection pipelines for assemblies focusing on ESKAPE pathogens in addition to Salmonella. This dataset consists of closed genomes from NCBI where paired-end Illumina data was available. The closed genomes from NCBI are provided in addition to assemblies based on the raw uploaded to NCBI, along with assemblies based on just the reads which mapped correctly to the closed assemblies. Reads were assembled using shovill v. 1.1.0 using SKESA and SPADES. Variants were called using snippy v. 4.6.0 and the mapped reads were extracted bedtools bamtofastq v2.29.2.</p>
AMR Benchmarking dataset - Mapped ReadSets - 5
<p>Benchmarking dataset for AMR detection pipelines for assemblies focusing on ESKAPE pathogens in addition to Salmonella. This dataset consists of closed genomes from NCBI where paired-end Illumina data was available. The closed genomes from NCBI are provided in addition to assemblies based on the raw uploaded to NCBI, along with assemblies based on just the reads which mapped correctly to the closed assemblies. Reads were assembled using shovill v. 1.1.0 using SKESA and SPADES. Variants were called using snippy v. 4.6.0 and the mapped reads were extracted bedtools bamtofastq v2.29.2.</p>
BasqueRoads, a dataset to benchmark road selection algorithms
<p>This dataset contains two road networks, one used for 1:25k scale maps (roads_ini), and one used at the 1:80k scale (roads_final). This dataset can be used to benchmark road selection algorithm that seek to transform the initial dataset into the final dataset.</p>
AMR Benchmarking dataset - Mapped ReadSets - 6
<p>Benchmarking dataset for AMR detection pipelines for assemblies focusing on ESKAPE pathogens in addition to Salmonella. This dataset consists of closed genomes from NCBI where paired-end Illumina data was available. The closed genomes from NCBI are provided in addition to assemblies based on the raw uploaded to NCBI, along with assemblies based on just the reads which mapped correctly to the closed assemblies. Reads were assembled using shovill v. 1.1.0 using SKESA and SPADES. Variants were called using snippy v. 4.6.0 and the mapped reads were extracted bedtools bamtofastq v2.29.2.</p>
maDLC Tri-Mouse Benchmark Dataset - Training
<p>see https://benchmark.deeplabcut.org/ for more information.</p>
maDLC Marmoset Benchmark Dataset - Training
<p>see https://benchmark.deeplabcut.org/ for more information.</p>
maDLC Fish Benchmark Dataset - Training
<p>see benchmark.deeplabcut.org for more information.</p>
Let's Trace It: Fine-Grained Serverless Benchmarking using Synchronous and Asynchronous Orchestrated Applications - Dataset
<p>This dataset contains the raw collected traces, preprocessed versions of the traces, and summary figures for the data associated with our manuscript <em>Let's Trace It: Fine-Grained Serverless Benchmarking using Synchronous and Asynchronous Orchestrated Applications.</em></p> <p>It contains over 7.5 million (7 564 830) traces of the ten applications integrated with ServiBench. The measurements were conducted on AWS Lambda in the us-east-1 region in late 2021 and early 2022. For more details on how the traces were collected we refer to our manuscript.</p> <p>For details on how to replicate our existing analysis on this dataset, we refer to https://github.com/ServiBench/ReplicationPackage</p>
MDMcleaner benchmarking datasets
<p>Benchmarking mock SAG datasets used for validation of the MDMcleaner workflow</p>
haRTStone - Benchmark Classification Datasets
<p><a href="https://www.tuhh.de/es/esd/research/projects/hartstone.html">Automated Generation of Benchmark Programs for the Evaluation of Analyses and Optimizations for Hard Real-Time Systems</a></p> <p>Many embedded systems are safety-critical real-time systems that have to meet hard deadlines (e.g., airbag or flight control systems). When designing such real-time systems, it is of utmost importance to guarantee that all tasks of a system meet their given deadlines. For this purpose, dedicated timing analyses are required that examine the worst-case behavior of a system and are able to provide such guarantees. In the case that deadlines are not met, optimizations need to be applied in order to modify the code of the system such that timing constraints are nevertheless finally met.</p> <p>Research on such analyses and optimizations for hard real-time systems is an extremely lively area where new results are presented regularly and at a very fast pace. Naturally, the evaluation of such analyses and optimizations plays a very important role. Nowadays, evaluation typically relies on benchmarking such that new analyses or optimizations are applied to existing collections of applications, tasks or program codes. The currently used benchmarks are, however, highly limited and not sufficient in order to perform a sound and scientific evaluation, especially if massively parallel multi-task systems are considered.</p> <p>For a well-founded and reproducible evaluation of analyses and optimizations, there is a strong demand for universally applicable benchmark approaches that are freely available for the entire scientific community. Benchmarks should satisfy the needs and requirements of various branches of research (e.g., schedulability analysis, WCET analysis, compiler optimization) on the one hand, but should also, on the other hand, realistically represent different application domains like, e.g., control or signal processing applications.</p> <p>This project aims at the realization of a flexible and parameterizable benchmark generator that produces benchmark programs in an automated, pseudo-randomized and reproducible fashion. This benchmark generator will in particular cover the system and the code level by producing both complete task sets and also actual program codes for the individual tasks. In order to enable a widespread use of the generator and a broad collaboration with arbitrary interested people and groups, this project will be inclusive and the developed software will be openly available right from the beginning. In the end, this project shall lead to a methodology for benchmarking-based evaluation that describes clearly and reproducibly for the different real-time communities, how to use the benchmark generator in order to obtain plausible, sound and scientifically accepted evaluation results.</p> <p>In order to be able to generate realistic and useful benchmarks, it is necessary to characterize key features of real-life applications and benchmarks, and to classify such applications according to their respective application domains. For this purpose, this archive contains datasets of the classification of existing ANSI-C benchmarks into their respective application domains, as well as trained DGCNN models of such classifiers.</p> <p>The freely accessible GitLab repository of the haRTStone project can be found at <a href="https://collaborating.tuhh.de/chf2198/haRTStone">https://collaborating.tuhh.de/chf2198/haRTStone</a>.</p>
A benchmarking dataset for peak detection methods in untargeted metabolomics using LC/HRMS
<p>Randomly selected 20,000 mz and rt pairs were manually evaluated for true positive and true negative signals in a single LC/HRMS data file (003.mzML) from the study MTBLS1684. The data file was processed using four peak detection methods - IDSL.IPA, XCMS, MZMINE and MSDIAL. </p>
Sample datasets for a tool benchmark
<p>Processed single cell datasets for a tool benchmark.</p> <p>Raw data was downloaded from GEO:</p> <p>1) https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE122960</p> <p>2) https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE128033</p> <p>3) https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE135893</p> <p>4) https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE136831</p> <p> </p> <p>Normalization procedure can be found here: https://github.com/mora-lab/cell-cell-interactions/tree/main/benchmark-workflow/R</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.