Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
369
datasets available to search
ShareScore release 0.7.1
Dataset results
369 results for “Benchmark Dataset”
A benchmark dataset of herbarium specimen images with label data: Summary
<p>This landing page contains a CSV file compiling all data associated with herbarium specimens that are part of this dataset, as they could be found on GBIF, JACQ or FinBIF. A CSV file with and without Darwin Core extension data is available, as some CSV readers have trouble with the JSON format that is used for those extensions.</p> <p>In addition, DOI's of the individual specimens uploaded to Zenodo and direct links to the different files (JPEG, TIFF, JSON, PNG) are also included. Index of these added variables:</p> <p>- persistentID: Persistent Identifier of the collection specimen. Data uploaded as part of this dataset will not be kept in sync with changes at the collection's repository. Hence, this URI will always point to the most up to date information known about the herbarium specimen.</p> <p>- jpegURL, tiffURL, jsonURL: URL's pointing straight to the respective image and data files themselves, to facilitate (selective) batch downloads.</p> <p>- pngSegAllURL and pngSegSelURL: Segmented overlays of the herbarium specimens indicating the location of different labels and reference material on the sheet ("All") and their content ("Sel"). More information can be found in the paper (in prep) associated with this data publication and the individual depositions themselves.</p> <p>- DOI: The DOI of the deposition of images and data of these specimens on Zenodo. DOI's point to the most up-to-date version of these depositions at the time of the publication of this CSV file. As a rule, this CSV file will be updated should any changes happen to any of the depositions.</p> <p>- jpegURL2, tiffURL2: A few herbarium sheets had labels on the back and consisted therefore of two scans. As a rule, the label scans are in this category.</p>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Annotation V2.0)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is a new version of the Annotation of MELA dataset, in which we add a missing annotation for 'mela_0732'. This file includes the annotations of the whole training set and validation set. </p> <p>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</p> <p> `public_id: anonymous patient ID to match images and annotations.<br> `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Validation Set and Annotation)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is the Validation Set and Annotation of MELA dataset, including 110 CTs and the annotations of the whole training set and validation set. Files include:</p> <ol> <li>Val.zip: 110 CTs in NII format (nii.gz).</li> <li>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</li> </ol> <p> `public_id: anonymous patient ID to match images and annotations.<br> `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>
The benchmark datasets for Multi-class Change Detection (MCD)
<p>Change detection (CD) provides a research basis for environmental monitoring, urban expansion and reconstruction as well as disaster assessment, by identifying the changes of ground objects in different time periods. Traditional CD focused on the binary change detection (BCD), focusing solely on the change and no-change regions. Due to the dynamic progress of earth observation satellite techniques, the spatial resolution of remote sensing images continues to increase, multi-class change detection (MCD) which can reflect more detailed land change has become a hot research direction in the field of CD.<strong> </strong></p> <p>We have collected the current open source benchmark datasets in the MCD of remote sensing imagery , in order to facilitate the sharing of the latest research datasets in the MCD field. Users can access the relevant MCD datasets through the links in the files.</p> <p>Source:</p> <p>Q. Zhu, X. Guo, Ziqi Li, D. Li*, “A review of Multi-class Change Detection for Remote Sensing Imagery” Geo-spatial information science, 2022</p>
Twist Whole-Exome Sequencing Dataset - High Coverage - WGGC SIG4 Benchmarking
<p>GIAB Reference Genome for Benchmarking Initiatives in the West German Genome Center (WGGC) - SIG4. </p> <p>Twist Whole-Exome Sequencing Dataset - High Coverage - 200M Reads.</p> <p> </p> <p> </p>
Unsupervised New Physics detection at 40 MHz: h+ -> tau nu Signal Benchmark Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of h+ -> tau nu decays produced in collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
Unsupervised New Physics detection at 40 MHz: h^0 -> tau tau Signal Benchmark Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of h^0 -> tau tau decays produced in collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
Unsupervised New Physics detection at 40 MHz: LQ -> b tau Signal Benchmark Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of Leptoquarks -> b tau decays produced in collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
Unsupervised New Physics detection at 40 MHz: A -> 4 leptons Signal Benchmark Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of A -> 4 leptons decays produced in collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
TCAB: Text Classification Attack Benchmark Dataset
<p>TCAB is a large collection of successful adversarial attacks on state-of-the-art text classification models trained on multiple sentiment and abuse domain datasets.</p> <p>The dataset is broken up into 2 files: <em>train.csv and</em> <em>val.csv</em>. The training set contains 1,448,751 instances (552,364 are "clean" unperturbed instances) and the validation set contains 482,914 instances (178,607 are "clean"). Each instance contains the following attributes:</p> <p><strong>scenario</strong>: Domain, either <em>abuse</em> or <em>sentiment</em>.</p> <p><strong>target_model_dataset</strong>: Dataset being attacked.</p> <p><strong>target_model_train_dataset</strong>: Dataset the target model trained on.</p> <p><strong>target_model</strong>: Type of victim model (e.g., <em>bert</em>, <em>roberta</em>, <em>xlnet</em>).</p> <p><strong>attack_toolchain</strong>: Open-source attack toolchain, either TextAttack or OpenAttack.</p> <p><strong>attack_name</strong>: Name of the attack method.</p> <p><strong>original_text</strong>: Original input text.</p> <p><strong>original_output</strong>: Prediction probabilities of the target model on the original text.</p> <p><strong>ground_truth</strong>: Encoded label for the original task of the domain dataset. 1 and 0 means toxic and toxic for abuse datasets, respectively. 1 and 0 means positive and negative sentiment for sentiment datasets. If there is a neutral sentiment, then 2, 1, 0 means positive, neutral, and negative sentiment.</p> <p><strong>status</strong>: Unperturbed example if "clean"; successful adversarial attack if "success".</p> <p><strong>perturbed_text</strong>: Text after it has been perturbed by an attack.</p> <p><strong>perturbed_output</strong>: Prediction probabilities of the target model on the perturbed text.</p> <p><strong>attack_time</strong>: Time taken to execute the attack.</p> <p><strong>num_queries</strong>: Number of queries performed while attacking.</p> <p><strong>frac_words_changed</strong>: Fraction of words changed due to an attack.</p> <p><strong>test_index</strong>: Index of each unique source example (original instance) (LEGACY - necessary for backwards compatibility).</p> <p><strong>original_text_identifier</strong>: Index of each unique source example (original instance).</p> <p><strong>unique_src_instance_identifier</strong>: Primary key to uniquely identify to every source instance; comprised of (<em>target_model_dataset</em>, <em>test_index</em>, <em>original_text_identifier</em>).</p> <p><strong>pk</strong>: Primary key to uniquely identify every attack instance; comprised of (<em>attack_name</em>, <em>attack_toolchain</em>, <em>original_text_identifier</em>, <em>scenario</em>, <em>target_model</em>, <em>target_model_dataset</em>, <em>test_index).</em></p>
Simulated Arabidopsis thaliana sequencing datasets for chloroplast assembler benchmarking
<p><strong>Changes</strong></p> <ul> <li>Fixed non-circular sampling from chloroplast and mitochondrion in version 1.1.0</li> <li>Fixed off-by-one error in reverse read in version 1.0.0</li> </ul> <p><strong>Purpose and Documentation</strong></p> <p>See: <a href="https://github.com/chloroExtractorTeam/benchmark">github.com/chloroExtractorTeam/benchmark</a></p> <p><strong>Original data</strong><br> The original <em>Arabidopsis thaliana </em>sequences were downloaded from TAIR: </p> <p>The Arabidopsis Information Resource (<a>TAIR</a>) on www.arabidopsis.org, Mar 22, 2019 available under the <a href="http://www.arabidopsis.org/doc/about/tair_terms_of_use/417">TAIR Terms of Use</a> </p> <p><em>Tanya Z. Berardini, Leonore Reiser, Donghui Li, Yarik Mezheritsky, Robert Muller, Emily Strait and Eva Huala. "The Arabidopsis Information Resource: Making and mining the "gold standard" annotated reference plant genome." genesis 2015 <a href="https://doi.org/10.1002/dvg.22877">doi:10.1002/dvg.22877</a></em></p> <p><strong>Programs used to generate this data</strong><br> - <a href="https://github.com/shenwei356/seqkit">seqkit</a> (v0.10.1): Shen W, Le S, Li Y, Hu F (2016) "SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA/Q File Manipulation." PLOS ONE 11(10): e0163962. <a href="https://doi.org/10.1371/journal.pone.0163962">doi:10.1371/journal.pone.0163962</a></p> <p> </p>
UnientrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers
<p>Our work focuses on providing a comprehensive dataset and benchmarks for evaluating gene ontology annotations using a unified system of Entrez Gene Identifiers.</p>
Datasets used in the benchmarking study of MR methods
<p>We conducted a benchmarking analysis of 16 summary-level data-based MR methods for causal inference with five real-world genetic datasets, focusing on three key aspects: type I error control, the accuracy of causal effect estimates, replicability, and power.</p> <p>The datasets used in the MR benchmarking study can be downloaded here:</p> <ol> <li>"dataset-GWASATLAS-negativecontrol.zip": the GWASATLAS dataset for evaluation of type I error control in confounding scenario (a): Population stratification</li> <li>"dataset-NealeLab-negativecontrol.zip": the Neale Lab dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-PanUKBB-negativecontrol.zip": the Pan UKBB dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-Pleiotropy-negativecontrol": the dataset used for evaluation of type I error control in confounding scenario (b): Pleiotropy;</li> <li>"dataset-familylevelconf-negativecontrol.zip": the dataset used for evaluation of type I error control in confounding scenario (c): Family-level confounders;</li> <li>"dataset_ukb-ukb.zip": the dataset used for evaluation of the accuracy of causal effect estimates;</li> <li>"dataset-LDL-CAD_clumped.zip": the dataset used for evaluation of replicability and power;</li> </ol> <p>Each of the datasets contains the following files:</p> <ol> <li> "Tested Trait pairs": the exposure-outcome trait pairs to be analyzed;</li> <li>"MRdat" refers to the summary statistics after performing IV selection (p-value < 5e-05) and PLINK LD clumping with a clumping window size of 1000kb and an r^2 threshold of 0.001.</li> <li>"bg_paras" are the estimated background parameters "Omega" and "C" which will be used for MR estimation in MR-APSS.</li> </ol> <p>Note:</p> <ol> <li>The formatted dataset after quality control can be accessible at our GitHub website (https://github.com/YangLabHKUST/MRbenchmarking).</li> <li>The details on quality control of GWAS summary statistics, formatting GWASs, and LD clumping for IV selection can be found on the MR-APSS software tutorial on the MR-APSS website (https://github.com/YangLabHKUST/MR-APSS).</li> <li>R code for running MR methods is also available at https://github.com/YangLabHKUST/MRbenchmarking.</li> </ol>
Dataset of a multiphase flow and reactive transport benchmark for radioactive waste disposal
<p>The files include the full dataset (tables and figures) of the comparion the results of a multiphase flow and reactive transport<br>benchmark for radioactive waste disposal. The codes INVERSE-FADES-CORE V2, DuMuX , TOUGHREACT and<br>iCP were benchmarked with 6 test cases of increasing complexity, starting with conservative tracer transport under variably<br>unsaturated conditions and ending with water flow, gas diffusion, minerals and cation exchange.</p>
Metadata dataset: benchmark datasets for modelling
<p>The main goal of the Soil Mission MARVIC project is to develop a framework for designing harmonized context-specific Monitoring, Reporting and Verification (MRV) systems for carbon farming, in support of the EU Carbon Removals and Carbon Farming (CRCF) regulation.</p> <p>The scope of this report (MARVIC Deliverable 2.1) is to provide a metadata dataset of benchmark sites (BS) that are relevant to the calibration and validation of models used within the MARVIC test cases. The dataset provided by Deliverable 2.1 describes the main characteristics of each site, such as pedoclimatic conditions, management practices applied, soil chemical, physical, and biological parameters, and details the measured variables that have been collected over time. </p> <p>This dataset of metadata is used within MARVIC to determine which modelling approaches can be used in each of the test cases across work packages. </p>
OpenMapCD: A Multimodal Benchmark Dataset for Change Detection Between Optical Remote Sensing and Map Data
<p><strong>Overview: </strong></p> <ol> <li>OpenMapCD, the <strong>first large-scale multimodal dataset</strong> for change detection on optical remote sensing imagery and map (OpenStreetMap) data, <strong>supporing basic binary change detection and further semantic change detection</strong></li> <li>OpenMapCD is highly geographically diverse, with <strong>1288</strong> benchmark samples with 1024x1024 pixels from <strong>40 </strong>regions across six continents and out-of-distribution data in two areas in Japan</li> <li>Advancing land-cover mapping, binary change detection and semantic change detection tasks, and GIS system updating<br><br></li> </ol> <p><strong>Research Paper: <br></strong></p> <ul> <li>Arxiv paper: <a href="https://arxiv.org/abs/2310.02674v3">https://arxiv.org/html/2310.02674v3</a></li> <li>TGRS paper: <a href="https://ieeexplore.ieee.org/document/10551264">https://ieeexplore.ieee.org/document/10551264</a></li> </ul> <p><strong><br>Project Page:</strong><br>The benchmark code is available at: <a href="https://github.com/ChenHongruixuan/ObjFormer">https://github.com/ChenHongruixuan/ObjFormer</a><br><br><strong>Reference:</strong></p> <pre><code>@ARTICLE{Chen2024ObjFormer, author={Chen, Hongruixuan and Lan, Cuiling and Song, Jian and Broni-Bediako, Clifford and Xia, Junshi and Yokoya, Naoto}, journal={IEEE Transactions on Geoscience and Remote Sensing}, title={ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer}, year={2024}, volume={62}, number={}, pages={1-22}, doi={10.1109/TGRS.2024.3410389} }</code></pre>
CLDF dataset derived from List and Prokić's "Benchmark Database of Phonetic Alignments" from 2014
<p>Cite the source of the dataset as:</p> <blockquote> <p>List, Johann-Mattis and Jelena Prokić. (2014). A benchmark database of phonetic alignments in historical linguistics and dialectology. In: Proceedings of the International Conference on Language Resources and Evaluation (LREC), 26 — 31 May 2014, Reykjavik. 288-294.</p> </blockquote>
Benchmark dataset for arby
<p>Datasets for the benchmarks performed on the reduce_basis function of the arby project <a href="https://arby.readthedocs.io/">https://arby.readthedocs.io/</a></p>
CrowdSpeech and Vox DIY: Benchmark Dataset for Crowdsourced Audio Transcription
<p>We collect and release CrowdSpeech — the first publicly available large-scale dataset of crowdsourced audio transcriptions. e show its applicability on an under-resourced language by constructing VoxDIY — a counterpart of CrowdSpeech for the Russian language.</p>
Materials Science Optimization Benchmark Dataset for Multi-Objective, Multi-Fidelity Optimization of Hard-Sphere Packing Simulations
<p>Benchmarks are an essential driver of progress in scientific disciplines. Ideal benchmarks mimic real-world tasks as closely as possible, where insufficient difficulty or applicability can stunt growth in the field. Benchmarks should also have sufficiently low computational overhead to promote accessibility and repeatability. The goal is then to win a “Turing test” of sorts by creating a surrogate model that is indistinguishable from the ground truth observation (at least within the dataset bounds that were explored), necessitating a large amount of data. In the fields of materials science and chemistry, industry-relevant optimization tasks are often hierarchical, noisy, multi-fidelity, multi-objective, high-dimensional, and non-linearly correlated while exhibiting mixed numerical and categorical variables subject to linear and non-linear constraints. To complicate matters, unexpected, failed simulation or experimental regions may be present in the search space. In this study, 494498 random hard-sphere packing simulations representing 206 CPU days worth of computational overhead were performed across nine input parameters with linear constraints and two discrete fidelities each with continuous fidelity parameters and results were logged to a free-tier shared MongoDB Atlas database. Two core tabular datasets resulted from this study: 1. a failure probability dataset containing unique input parameter sets and the estimated probabilities that the simulation will fail at each of the two steps, and 2. a regression dataset mapping input parameter sets (including repeats) to particle packing fractions and computational runtimes for each of the two steps. These two datasets are used to create a surrogate model as close as possible to running the actual simulations by incorporating simulation failure and heteroskedastic noise. For the regression dataset, percentile ranks were computed within each of the groups of identical parameter sets to enable capturing heteroskedastic noise. This is in contrast with a more traditional approach that imposes a-priori assumptions such as Gaussian noise e.g., by providing a mean and standard deviation. A similar approach can be applied to other benchmark datasets to bridge the gap between optimization benchmarks with low computational overhead and realistically complex, real-world optimization scenarios.</p> <p>For usage instructions, see https://matsci-opt-benchmarks.readthedocs.io/.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.