Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

369

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

369 results for “Datasets Benchmarking”

Learn how ShareScore rates datasets ↗
zenodo36/100

AMR Benchmarking dataset - Metagenomics

<p>Metagenomic benchmarking dataset for AMR detection pipelines for assemblies focusing on ESKAPE pathogens in addition to Salmonella. This dataset consists of closed genomes from NCBI where paired-end Illumina data was available. These genomes were then randomly assigned a relative abundance, had additional AMR genes randomly inserted (to cover all AMR genes in CARD v3.1.4) and metagenomic Illumina reads simulated from them.<br> <br> Metagenomic simulation was performed using: https://github.com/fmaguire/AMR_Metagenome_Simulator and the entire process can be repeated using the <a href="https://zenodo.org/api/files/ba12e743-ae3d-43c0-98bd-2075d6f50340/metagenome_benchmark.sh">metagenome_benchmark.sh </a>script included above.<br> &nbsp;</p> <p><strong>Files</strong><br> `amr_benchmarking_metagenome.csv` contains the metadata the input genome accessions, paths, and simulated copy number used for creation of the AMR metagenome.</p> <p>`AMR_metagenome_labels.tsv` a two column csv containing names of all reads that are derived from an AMR gene and an identifier for the corresponding AMR gene. AMR genes are identified using CARD Antibiotic Resistance Ontology (ARO), with a suffix listing any SNVs for nmutation related resistance genes.<br> &nbsp;</p> <p>`simulated_metagenome.fna.gz` contains the full &quot;assembled&quot; true metagenomic contigs (derived directly from the input genome assemblies amplified to the correct copy number).</p> <p>`metagenome_unsorted.bed` contains the location of AMR genes in the full &quot;assembled&quot; true metagenomic contigs<br> <br> `simulated_metagenome_{1,2}.fq.gz` contain the simulated paired end metagenomics reads</p> <p>`simulated_metagenome_error_free.bam` contains the error-free mapping location from which simulated reads were derived</p> <table> <tbody> <tr> <td>&nbsp;</td> </tr> <tr> <td>&nbsp;</td> </tr> </tbody> </table>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Bangla Text Normalization Benchmark Dataset

<p>This is a small benchmark dataset for Bangla Text Normalization.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

dataset How does the balance of metal and acid functions on the benchmark Mo/ZSM-5 catalyst drive the Methane dehydroaromatization reaction

<p>dataset How does the balance of metal and acid functions on the benchmark Mo/ZSM-5 catalyst drive the Methane dehydroaromatization reaction</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Dataset for "A randomized benchmarking suite for mid-circuit measurements"

<p>Dataset and simulation code to generate the results of&nbsp;&quot;A randomized benchmarking suite for mid-circuit measurements&quot;.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

A benchmark dataset for the flood event of Mandra (Athens, Greece), 2017

<p>This dataset includes:</p> <p>1) The maximum water depths recorded at the 44 points after the flood event which hit Mandra, Athens, Greece (15th of November 2017) with their coordinates in both GGRS87 (Greek Geodetic System) and WGS84 systems. Besides, the simulated maximum water depths as derived by the HEC-RAS software (which is calibrated) and the MIKE FlOOD software (informed by the calibrated HEC-RAS).</p> <p>2) The Digital Terrain Model (DTM) of the greater area with a resolution of 5 m, as provided by the National Cadastre and Mapping Agency of Greece.</p> <p>3) Two Shape files with (a) the computational area and the (b) upsream/boundary conditions used for both calibrated HEC-RAS and informed MIKE FLOOD software.</p> <p>4) A shape file with the Mandra urban blocks footprint, as were manually drawn on Google Earth platform.</p> <p>5) The ensemble of 100 hydrographs which serve as the inflow from Agia Aikaterini catchment to the Mandra town. They are derived by implementing the FLOW-R2D hydrodynamic simulator at the catchment scale, having as an input the rainfall field captured by the weather radar during the Mandra flood event (Bellos et al., 2020). &nbsp;</p> <p>6) The rainfall field of the greater area with a spatial resolution of 200 m and a temporal resolution of 2 min, recorded by the X-band weather radar of the National Observatory of Athens (Bellos et al., 2020).</p> <p><br> References:</p> <p>Bellos, V., Papageorgaki, I., Kourtis, I., Vangelis, H., Kalogiros, I., Tsakiris, G. (2020). Reconstruction of a flash flood event using a 2D hydrodynamic model under spatial and temporal variability of storm. Natural Hazards, 101(3), 711-726.</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

SMiCRM: A Benchmark Dataset of Mechanistic Molecular Images

<p>Optical chemical structure recognition (OCSR) systems aim to extract the molecular structure information, usually in the form of molecular graph or SMILES, from images of chemical molecules. While many tools have been developed for this purpose, challenges still exist due to different types of noises that might exist in the images. Specifically, we focus on the &ldquo;arrow-pushing&rdquo; diagrams, a typical type of chemical images to demonstrate electron flow in mechanistic steps. We present Structural molecular identifier of Molecular images in Chemical Reaction Mechanisms (SMiCRM), a dataset designed to benchmark machine recognition capabilities of chemical molecules with arrow-pushing annotations. Comprising 453 images, it spans a broad array of organic chemical reactions, each illustrated with molecular structures and mechanistic arrows. SMiCRM offers a rich collection of annotated molecule images for enhancing the benchmarking process for OCSR methods. This dataset includes a machine-readable molecular identity for each image as well as mechanistic arrows showing electron flow during chemical reactions. It presents a more authentic and challenging task for testing molecular recognition technologies, and achieving this task can greatly enrich the mechanisitic information in computer-extracted chemical reaction data.</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Benchmark datasets for testing AIRR-seq data processing with pyIgMap pipeline

<p>This is a set of raw FASTQ files produced by various AIRR-seq protocols, stored here for benchmark conveinience and for reference purposes.</p> <p>Currently it contains the following datasets:</p> <table> <tbody> <tr> <td><strong>id</strong></td> <td><strong>fastq</strong></td> <td><strong>reference</strong></td> <td><strong>method</strong></td> <td><strong>description</strong></td> </tr> <tr> <td>allergy</td> <td> <p>ERR7425614_1.fastq.gz,</p> <p>ERR7425614_2.fastq.gz</p> </td> <td>https://doi.org/10.7554/eLife.79254</td> <td>5'RACE, UMI, long MiSeq reads, IGH w/ isotype</td> <td>Longitudinal full-length IGH repertoire profiling and clonal lineage dynamics in memory B cells, plasmablasts and plasma cells of human peripheral blood</td> </tr> <tr> <td>covid</td> <td> <p>fmba_TRAB_R1.fastq.gz,</p> <p>fmba_TRAB_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1101/2023.11.08.566227</p> </td> <td>DNA multiplex, UMI, TRA+TRB mix, NextSeq</td> <td>TCR sequencing in COVID-19 convalescent and healthy donors</td> </tr> <tr> <td>natprot</td> <td> <p>PMID27490633_R1.fastq.gz,</p> <p>PMID27490633_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1038/nprot.2016.093</p> </td> <td>5'RACE, UMI, long MiSeq reads, high-quality overlap, IGH no isotype</td> <td>High-quality full-length immunoglobulin profiling with unique molecular barcoding</td> </tr> <tr> <td>brnaseq</td> <td> <p>SRR3743469_R1.fastq.gz,</p> <p>SRR3743469_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1016/j.immuni.2016.08.012</p> </td> <td>RNA-Seq, all chains, B-cells</td> <td>Primary human mature na&iuml;ve B-cells (IgD+CD38lo; NB) and GCB-cells (CD77+CD38hi; GCB) were purified from tonsils of healthy individuals. RNA-seq libraries were prepared using the Illumina TruSeq RNA sample kits according to the manufacturer.</td> </tr> <tr> <td>uhrr</td> <td> <p>UHRR_full_R1.fastq.gz,</p> <p>UHRR_full_R2.fastq.gz</p> </td> <td> <p>https://doi.org/10.1038/s41598-021-04583-z</p> </td> <td>RNA-Seq, bulk</td> <td>Universal Human Reference RNA</td> </tr> <tr> <td>10x</td> <td> <p>10x_bcr_R1.fastq.gz,</p> <p>10x_bcr_R2.fastq.gz,</p> <p>10x_tcr_R1.fastq.gz,</p> <p>10x_tcr_R2.fastq.gz</p> </td> <td> <p>https://www.10xgenomics.com/datasets/human-pbmc-from-a-healthy-donor-10-k-cells-v-2-2-standard-5-0-0</p> </td> <td>10x Genomics vdj</td> <td>See reference. AIRR-seq is split into TCR and BCR parts</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Large-Scale Multipurpose Benchmark Datasets For Assessing Data-Driven Deep Learning Approaches For Water Distribution Networks

<p>&nbsp;</p> <div> <div><a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Tello,+A">Andres Tello*</a><em>, </em><a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Truong,+H">Huy Truong*</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Lazovik,+A">Alexander Lazovik</a>, <a href="https://arxiv.org/search/cs?searchtype=author&amp;query=Degeler,+V">Victoria Degeler</a>. Large-Scale Multipurpose Benchmark Datasets For Assessing Data-Driven Deep Learning Approaches For Water Distribution Networks. Engineering Proceedings. 2024; 69(1):50. <a href="https://doi.org/10.3390/engproc2024069050">https://doi.org/10.3390/engproc2024069050</a></div> <br> <div>(*) Both authors contributed equally.<br><br></div> <h2>Update</h2> <div>(04/09/2024): Citation is updated.<br>We have added headers for CSVs and auxiliary data (duration time, edge list, ordered names.. ) in the configuration file (JSON format). As such, corresponding INP files can be omitted when working with this version.&nbsp;<br>The EXN network has been included in this version, so the total number of processed networks is 11.<br>For more details, please read ZENODO_README.md.</div> <h2>Contact</h2> <div>For dataset-related questions: <a href="mailto:h.c.truong@rug.nl" target="_blank" rel="noopener">Huy Truong</a></div> <br> <div>For data acquisition: <a href="mailto:a.tello@rug.nl" target="_blank" rel="noopener">Andres Tello</a></div> <br> <div>If you use this dataset, please cite:</div> <blockquote>@article{tello2024largescale,<br>&nbsp; &nbsp; AUTHOR = {Tello, Andr&eacute;s and Truong, Huy and Lazovik, Alexander and Degeler, Victoria},<br>&nbsp; &nbsp; TITLE = {Large-Scale Multipurpose Benchmark Datasets for Assessing Data-Driven Deep Learning Approaches for Water Distribution Networks},<br>&nbsp; &nbsp; JOURNAL = {Engineering Proceedings},<br>&nbsp; &nbsp; VOLUME = {69},<br>&nbsp; &nbsp; YEAR = {2024},<br>&nbsp; &nbsp; NUMBER = {1},<br>&nbsp; &nbsp; ARTICLE-NUMBER = {50},<br>&nbsp; &nbsp; URL = {https://www.mdpi.com/2673-4591/69/1/50},<br>&nbsp; &nbsp; ISSN = {2673-4591},<br>&nbsp; &nbsp; DOI = {10.3390/engproc2024069050}<br>}</blockquote> </div>

opencc-by-4.0May 2024View details →
zenodo36/100

Dataset Retrieval Benchmark

<p>The sample of dataset retrieval benchmark corpus:</p> <ul> <li><strong>docs.tsv</strong> contains 9802 dataset metadata records</li> <li><strong>topics.tsv, keyword_queries.tsv</strong> contain 51 original post and keywords query for each post</li> <li><strong>qrels.tsv</strong> contains 57 relevance judgements for provided sample of queries</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Benchmark dataset for CATH hierarchical clustering tools (GeMMA/FunFHMMEr, MARC, FRAN and eMMA)

<p>Benchmark dataset for CATH SuperFamily 3.40.50.620 (HUPS).</p> <p>Contains Functional Families alignments and Hidden Markov Models generated by GeMMA/FunFHMMER, MARC, FRAN and CATH-eMMA and Python code used to assess their quality (EC purity, DOPS, Neff) and intermediate steps by the MARC and FRAN pipelines (pooling, randomisation, renaming).</p> <p>3.4.50.620_full_superfamily_sequences.fasta contains all HUPs superfamily sequences, the FunFams are a subset of these.</p> <p>all_starting_clusters_sequences.fasta contain the sequences included in the starting clusters used in the analyses.</p> <p>3.40.50.620_embedded.pt includes embeddings for the HUPs superfamily generated using the ESM2 Protein Language Model.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

ProcarySV Artificial Benchmarking Datasets

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo36/100

Benchmark cancer datasets for Clustering algorithms for Omics-based Patient Stratification (COPS)

<p>This repository contains seven multi-omic cancer datasets including several cancer types (breast, kidney, lung, ovary, prostate, and thyroid cancers as well as low grade gliomas) that were used for benchmarking several multi-view clustering algorithms implemented by COPS (https://github.com/UEFBiomedicalInformaticsLab/COPS). The datasets were originally compiled from The Cancer Genoma Atlas (TCGA) and downloaded using the <em>curatedTCGAData</em> R-package. The datasets include copy-number variations, methylomics as well as mRNA and miRNA transcriptomics. The methylomics data was mapped to genes by averaging methylation level of probes associated with the promoter regions of genes. Similarly the miRNA transcriptomics data was mapped to genes by using known and predicted miRNA -&gt; gene interactions. Updated survival data was acquired from the Liu et al. 2018 paper.&nbsp;</p> <p>This repository also includes two sets of cancer associated pathway networks used by pathway-based multi-omic methods benchmarked in our study. NCI-PID pathways were downloaded using the <em>ndexr</em> R-package on December 22 2021. While KEGG pathways were downloaded using the <em>pathview</em> R-package on May 3 2022.&nbsp;</p> <p>More details on the processing can be found on the related publication.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

SEERS dataset for benchmarking

<p>This&nbsp;dataset&nbsp;contains&nbsp;data collected during the demonstration activities of the SEERS project (http://www.seersproject.eu) including four scenarios: firefighting in enclosed spaces, maritime aerial surveillance, port surveillance and traffic monitoring.&nbsp;The purpose of the&nbsp;dataset is&nbsp;the evaluation of the video analytics software.&nbsp;</p>

opencc-by-4.0Feb 2018View details →
zenodo36/100

Heteroplasmy Benchmark Dataset - mitochondrial DNA mixture model - HiSeq - M1-M4 - BAM

<p>Illumina HiSeq data of mixtures M1 (50%), M2 (10%), M3 (2%) and M4 (1%) of haplotypes&nbsp;H1c6 and U5a2e (decreasing).</p> <p>See&nbsp;<a href="https://doi.org/10.1371/journal.pone.0135643">https://doi.org/10.1371/journal.pone.0135643</a>&nbsp;for technical/lab-related informations</p>

opencc-by-4.0Dec 2018View details →
zenodo36/100

PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano Performance

<p>Recently, artificial intelligence techniques for education have been received increasing attentions, while it still remains an open problem to design the effective music instrument instructing systems. Although key presses can be directly derived from sheet music, the transitional movements among key presses require more extensive guidance in piano performance. In this work, we construct a piano-hand motion generation benchmark to guide hand movements and fingerings for piano playing. To this end, we collect an annotated dataset, PianoMotion10M, consisting of 116 hours of piano playing videos from a bird's-eye view with 10 million annotated hand poses. We also introduce a powerful baseline model that generates hand motions from piano audios through a position predictor and a position-guided gesture generator. Furthermore, a series of evaluation metrics are designed to assess the performance of the baseline model, including motion similarity, smoothness, positional accuracy of left and right hands, and overall fidelity of movement distribution. Despite that piano key presses with respect to music scores or audios are already accessible, PianoMotion10M aims to provide guidance on piano fingering for instruction purposes.</p>

opencc-by-nc-nd-4.0May 2024View details →
zenodo36/100

DIVERSE: Deciphering Internet Views on the U.S. Military Through Video Comment Stance Analysis: A Novel Benchmark Dataset for Stance Classification

<p>Paper citation: Cruickshank, Iain J., and Lynnette Hui Xian Ng. "DIVERSE: Deciphering Internet Views on the US Military Through Video Comment Stance Analysis, A Novel Benchmark Dataset for Stance Classification." <em>arXiv preprint arXiv:2403.03334</em> (2024).</p> <p>Link to paper: https://arxiv.org/abs/2403.03334</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Plot-Rice v1.0: A global plot-based rice benchmark dataset with spatiotemporal heterogeneity for scientific deep learning

<p>This dataset (Plot-Rice v1.0) offers a global rice benchmark dataset for scientific deep learning at a 10-meter resolution for the year 2023. Plot-Rice v1.0 is constructed based on Sentinel-1 and Sentinel-2 images, encompassing plot-level rice labels and corresponding multi-source feature time series from 20 countries worldwide. It fully considers the spatiotemporal heterogeneity of rice and supports continuous updates, providing a data benchmark for performance comparisons of deep learning models.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Molecular surface coverage standards by reference-free GIXRF supporting SERS and SEIRA substrate benchmarking - Dataset

<p>This is the dataset of "Molecular surface coverage standards by reference-free GIXRF supporting SERS and SEIRA substrate benchmarking".&nbsp;</p> <p><a href="https://doi.org/10.1515/nanoph-2024-0222" target="_blank" rel="noopener">https://doi.org/10.1515/nanoph-2024-0222</a></p> <p>Part of this work was supported by the European project OpMetBat, code 21GRD01. The project has received funding from the European Partnership on Metrology, cofinanced from the the European Union's Horizon Europe Research and Innovation Programme, and by Participating States.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

NanoBaseLib: A Multi-Task Benchmark Dataset for Nanopore Sequencing

<p>NanoBaseLib is a multi-task benchmark dataset for Nanopore Sequencing. We compile and preprocess publicly available datasets using a unified pipeline to ensure consistency and quality across all tasks. The dataset is benchmarked for four key Nanopore sequencing tasks: base calling, polyA detection, segmentation and event alignment, and RNA modification detection. &nbsp;NanoBaseLib is available at <a href="https://nanobaselib.github.io/">https://nanobaselib.github.io</a>.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Expert Finding Benchmark Datasets (IR, CL and SW communities)

<p>This&nbsp;is the updated version&nbsp;of the original benchmark expert finding&nbsp;datasets proposed by the authors of this paper -&nbsp;<a href="https://doi.org/10.1145/2508497.2508501">https://doi.org/10.1145/2508497.2508501</a>. The current version is released as part of Neural Expert Finder (NEF), a novel expert finding approach utilizing transformer based pre-trained language models (in .csv format).</p>

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record