Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,549
datasets available to search
ShareScore release 0.9.0
Dataset results
1,549 results for “benchmarks”
Code to generate figures 3 and 4 of: "A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics."
<p>Code to generate figures 3 and 4 of the manuscript titled "A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics."</p> <p> </p>
Supplementary Material for 'Benchmarking Explanatory Models for Inertia Forecasting using Public Data of the Nordic Area'
<p>This data set is supplementary material for the paper 'Benchmarking Explanatory Models for Inertia Forecasting using Public Data of the Nordic Area' by Jemima Sophie Graham, Evelyn Heylen, and Fei Teng.</p> <p>This data set is intended for day-ahead inertia forecasting in the Nordic (Eastern Denmark, Finland, Norway, Sweden). It contains hourly data for the inertial energy (MVAs), day-ahead national demand forecast (MW), day-ahead wind power forecast (MW), day-ahead solar power forecast (MW), and interconnection flow (MW) in the Nordic between January 2016 and August 2020. </p>
Benchmarking eliminative radiomic feature selection for head and neck lymph node classification - Supplemental data
<p>Supplementary files for the publication "Benchmarking eliminative radiomic feature selection for head and neck lymph node classification"</p>
Expert Finding Benchmark Datasets (IR, CL and SW communities)
<p>This is the updated version of the original benchmark expert finding datasets proposed by the authors of this paper - <a href="https://doi.org/10.1145/2508497.2508501">https://doi.org/10.1145/2508497.2508501</a>. The current version is released as part of Neural Expert Finder (NEF), a novel expert finding approach utilizing transformer based pre-trained language models.</p>
ClevrTex: A Texture-Rich Benchmark for Unsupervised Multi-Object Segmentation
<p>There has been a recent surge in methods that aim to decompose and segment scenes into multiple objects in an unsupervised manner, i.e., unsupervised multi-object segmentation. Performing such a task is a long-standing goal of computer vision, offering to unlock object-level reasoning without requiring dense annotations to train segmentation models. Despite significant progress, current models are developed and trained on visually simple scenes depicting mono-colored objects on plain backgrounds. The natural world, however, is visually complex with confounding aspects such as diverse textures and complicated lighting effects. In this study, we present a new benchmark called ClevrTex, designed as the next challenge to compare, evaluate and analyze algorithms. ClevrTex features synthetic scenes with diverse shapes, textures and photo-mapped materials, created using physically based rendering techniques. ClevrTex has 50k examples depicting 3-10 objects arranged on a background, created using a catalog of 60 materials, and a further test set featuring 10k images created using 25 different materials. We benchmark a large set of recent unsupervised multi-object segmentation models on ClevrTex and find all state-of-the-art approaches fail to learn good representations in the textured setting, despite impressive performance on simpler data. We also create variants of the ClevrTex dataset, controlling for different aspects of scene complexity, and probe current approaches for individual shortcomings.</p> <p>Project webpage: https://www.robots.ox.ac.uk/~vgg/data/clevrtex/</p> <p>These are <strong>the variant datasets</strong>. Please see project page for links to the main dataset.</p>
ClevrTex: A Texture-Rich Benchmark for Unsupervised Multi-Object Segmentation
<p>There has been a recent surge in methods that aim to decompose and segment scenes into multiple objects in an unsupervised manner, i.e., unsupervised multi-object segmentation. Performing such a task is a long-standing goal of computer vision, offering to unlock object-level reasoning without requiring dense annotations to train segmentation models. Despite significant progress, current models are developed and trained on visually simple scenes depicting mono-colored objects on plain backgrounds. The natural world, however, is visually complex with confounding aspects such as diverse textures and complicated lighting effects. In this study, we present a new benchmark called ClevrTex, designed as the next challenge to compare, evaluate and analyze algorithms. ClevrTex features synthetic scenes with diverse shapes, textures and photo-mapped materials, created using physically based rendering techniques. ClevrTex has 50k examples depicting 3-10 objects arranged on a background, created using a catalog of 60 materials, and a further test set featuring 10k images created using 25 different materials. We benchmark a large set of recent unsupervised multi-object segmentation models on ClevrTex and find all state-of-the-art approaches fail to learn good representations in the textured setting, despite impressive performance on simpler data. We also create variants of the ClevrTex dataset, controlling for different aspects of scene complexity, and probe current approaches for individual shortcomings.</p> <p>Project webpage: https://www.robots.ox.ac.uk/~vgg/data/clevrtex/</p> <p>This is the <strong>main dataset and OOD test set. </strong>Please see project page for links to the dataset variants.</p>
Benchmarks for ApproxCov and ApproxMaxCov algorithms evaluation
<p>This deposit contains the benchmarks used for the evaluation of ApproxCov and ApproxMaxCov algorithms and their extensions. Implementations of the algorithms can be found at https://github.com/meelgroup/approxcov.<br> The folder two_values/cnf contains constraints of configurable systems. It is a subset of benchmarks from https://zenodo.org/record/4022395 and https://zenodo.org/record/3793090. In this set of benchmarks all features can have two values.<br> The folder two_values/samples contains sets of configurations computed with baital, quicksampler, and waps tools.<br> The folder mult_values contains samples and constraints of configurable systems where features can have finite number of values. These benchmarks are taken from the evaluation materials of the paper [1] and are converted to the input format of ApproxCov and ApproxMaxCov algorithms.</p> <p>[1] Brady J Garvin, Myra B Cohen, and Matthew B Dwyer. 2009. An improved meta-heuristic search for constrained interaction testing. In 2009 1st International Symposium on Search Based Software Engineering. IEEE, 13–22.</p>
Data for the NeonTreeEvaluation Benchmark
<p>This dataset is the data files for the NeonTreeEvaluation Benchmark for individual tree detection from airborne imagery. For each geographic site, given by the NEON four letter code (e.g HARV -> Harvard Forest), there are up to 4 files: a RGB image, a LiDAR tile, and a 426 band hyperpspectral file, and a 1m canopy height file. For more information on the benchmark, and the corresponding R package, see <a href="https://github.com/weecology/NeonTreeEvaluation_package">https://github.com/weecology/NeonTreeEvaluation_package</a> </p> <p>Training.zip and Evaluation.zip both have the same folder structure with RGB, Hyperspectral, LiDAR and CHM folders. All annotations are in the annotation.zip. Not all files in the evaluation.zip have corresponding annotations. We have included these to allow users to test their approaches, annotate new data, or use semi-supervised methods.</p>
Benchmark problems for transcranial ultrasound simulation: Datasets for intercomparison of compressional wave models
<p>This dataset contains the skull maps and modeling results associated with the forthcoming publication "Benchmark problems for transcranial ultrasound simulation: Intercomparison of compressional wave models".</p>
Data and Analysis for "On the Reliability of Coverage-based Fuzzer Benchmarking"
<pre><strong>Data and Analysis for "On the Reliability of Coverage-based Fuzzer Benchmarking"</strong> <strong>## Cite</strong> </pre> <pre><code>@inproceedings{benchmarking, author = {B{\"o}hme, Marcel and Szekeres, L{\'a}szl{\'o} and Metzman, Jonathan}, title = {On the Reliability of Coverage-based Fuzzer Benchmarking}, year = {2022}, booktitle = {Proceedings of the 44th International Conference on Software Engineering}, series = {ICSE '22}, pages = {1-13}, doi = {10.1145/3510003.3510230} }</code></pre> <pre> <strong>## Data Analysis</strong> The Jupyter notebook generating all tables and figures can be found in fuzzbench.manual.ipynb <strong>## Generated Images and Tables</strong> The generated data analysis artifacts are also available in this artifact. <strong>## Data</strong> All the data is available in the FuzzBench Reports and will be automatically downloaded. * 20 trials of 23 hours with 15 programs and 10 fuzzers. * Experiment name: 2021-02-17-bug-paper * Report: https://www.fuzzbench.com/reports/2021-02-17-bug-paper/index.html * Data: https://www.fuzzbench.com/reports/2021-02-17-bug-paper/data.csv.gz * Fuzzbench Commit: [38e344fef2f1079579391a0d9dcb52319f7051f2](https://github.com/google/fuzzbench/commits/38e344fef2f1079579391a0d9dcb52319f7051f2) * 30 trials of 23 hours with 11 programs and 10 fuzzers. * Experiment name: 2021-08-19-crash-s * Report: https://www.fuzzbench.com/reports/2021-08-19-crash-s/index.html and * Data: https://www.fuzzbench.com/reports/2021-08-19-crash-s/data.csv.gz * Fuzzbench Commit: db192b60815ac87f69ee0f7f3e37aeac71949e1b * 30 trials of 23 hours with 11 programs and 10 fuzzers. * Experiment name: 2021-08-19-crash-s2 * Report: https://www.fuzzbench.com/reports/2021-08-19-crash-s2/index.html and * Data: https://www.fuzzbench.com/reports/2021-08-19-crash-s2/data.csv.gz * Fuzzbench Commit: db192b60815ac87f69ee0f7f3e37aeac71949e1b The deduplicated data can be found in * 2021-02-17-bug-paper-fixed2.csv.gz * 2021-08-19-crash-s-fixed2.csv.gz * 2021-08-19-crash-s2-fixed2.csv.gz <strong>## Reproducibility</strong> </pre> <pre><code class="language-bash"># Download the precise version of FuzzBench used for the experiment git clone https://github.com/google/fuzzbench.git cd fuzzbench git checkout <Fuzzbench Commit> # Download the internal config file. curl https://storage.googleapis.com/[experiment-name]/config/experiment.yaml > /tmp/experiment-config.yaml make install-dependencies # Launch the experiment using paramters from the internal config file. PYTHONPATH=. python experiment/reproduce_experiment.py -c /tmp/experiment-config.yaml -e <new_experiment_name></code></pre> <p> </p>
Benchmark SIFET 2022 - Dataset 2: L'area urbana di Santa Marta
<p>Il secondo test field è stato individuato nella zona di Calle Larga Santa Marta e delle calli ad essa trasversali. La zona di Santa Marta, un'area residenziale di Venezia di recente edificazione, si caratterizza per la presenza di edifici di altezza compresa tra i 10 m e i 15 m circa, tra loro non molto distanti. La disposizione di questi oggetti architettonici fa si che l'area assuma la configurazione di un corridoio urbano.</p> <p>Tale area si presta dunque a valutare la possibilità di utilizzo di una tecnica di rilievo alternativa alla fotogrammetria tradizionale, poiché di difficile applicazione in questo contesto, e al rilievo laser scanning, in quanto i dati acquisiti sulla parte sommitale degli edifici risultano molto scarsi. </p> <p>Sono state effettuate due strisciate fotogrammetriche ponendo le camere ad altezze differenti, 2 m per il Set 1 e 4,5 m per il Set 2, mantenendo i medesimi punti di presa, in questo modo è stata acquisita una quantità di dati sufficiente anche per la parte alta dei fronti degli edifici. L'utilizzo delle camere sferiche risulta essere particolarmente vantaggioso per l'acquisizione di immagini ad altezze differenti poiché molto leggere e facilmente installabili su di un'asta telescopica .</p> <p>Durante questo test le camere sferiche sono state orientate mantenendo gli assi ottici ortogonali alle facciate degli edifici; in corrispondenza degli slarghi presenti lungo la calle e ogniqualvolta fosse necessario compiere una rotazione di 90° è stato adottato uno schema di presa cruciforme, con riferimento all'andamento principale. Per il Set 1 sono state acquisite 113 immagini con la camera GoPro MAX 360 e 114 immagini con le camere Nikon KeyMission 360 e Ricoh Theta Z1; per il Set 2 sono state acquisite 106 immagini con la camera GoPro MAX 360, 105 immagini con le camere Nikon KeyMission 360 e 104 immagini con la camera Ricoh Theta Z1.</p>
Benchmark SIFET 2022 - Dataset 1: Il Chiostro dei Tolentini
<p>Il primo test field è rappresentato dal portico del Chiostro del complesso dei Tolentini, sede dell'Università IUAV di Venezia. La presenza di elementi come gli archi e le volte a crociera definisce il portico come un oggetto architettonico chiuso, di cui ogni parte è potenzialmente oggetto di rilievo, sia le pareti che la copertura.</p> <p>Il caso studio si presta a testare la possibilità di utilizzo di una metodologia operativa che consente una notevole riduzione dei tempi di acquisizione. Nel caso della fotogrammetria con camere sferiche è sufficiente una sola strisciata per acquisire tutte le informazioni necessarie alla generazione del modello fotogrammetrico del porticato; per avere le stesse informazioni utilizzando la fotogrammetria tradizionale si sarebbe invece reso necessario effettuare un numero molto maggiore di strisciate.</p> <p>Durante questo test le camere sferiche sono state orientate mantenendo l'asse ottico ortogonale alla parete in mattoni e al porticato: é stata effettuata una presa per ogni campata dei quattro lati mentre per le campate angolari sono state acquisite due immagini, in posizioni differenti, l'una ruotata di 90° rispetto all'altra. Con ogni camera sono state quindi acquisite 36 immagini.</p>
Benchmark SIFET 2022 - Dataset 3: Il Rio de S. Barnaba
<p>Il terzo test field individuato è stato il Rio de San Barnaba. Il Rio si caratterizza per avere lunghi tratti in cui gli edifici affacciano direttamente su un canale di larghezza variabile tra i 5 m e i 10 m. Anche in questo caso la disposizione degli edifici fa sì che si l'area assuma la configurazione di un corridoio urbano, la cui caratteristica fondamentale è la possibilità di essere percorso esclusivamente utilizzando una barca.</p> <p>Valutare la possibilità di utilizzo delle camere sferiche per il rilievo dei rii veneziani è un tema di reale interesse poiché la configurazione spaziale di tali luoghi non permette l'utilizzo delle tradizionali tecniche di rilievo: la fotogrammetria tradizionale risulta essere di difficile applicazione in questo contesto, e il rilievo laser scanning diventa una tecnica inutilizzabile non disponendo di una base stabile di appoggio. </p> <p>Durante questo test le camere sferiche sono state orientate mantenendo gli assi ottici ortogonali alle facciate degli edifici per il Set 1 e ruotandole di 90°, ovvero paralleli alle facciate degli edifici, per il Set 2. Per il Set 1 sono state acquisite 109 immagini (62 dell'area riportata nella sezione "Schema delle prese" e 47 delle zone immediatamente vicine, di completamento) con la camera Nikon KeyMission 360 e 110 immagini (63 dell'area riportata nella sezione "Schema delle prese" e 47 delle zone immediatamente vicine, di completamento) con le camere GoPro MAX 360 e Ricoh Theta Z1; per il Set 2 sono state acquisite 57 immagini con la camera Nikon KeyMission 360 e 56 immagini con le camere GoPro MAX 360 e Ricoh Theta Z1.</p>
UniFrac Benchmarking Data 0.20.1.1
<p>UniFrac benchmark data, comparing <a href="https://github.com/biocore/unifrac/releases/tag/0.20.1.1">version 0.20.1.1</a><a href="https://github.com/biocore/unifrac/tree/d0656987ecfad9d9d28e7f95e9858d99e08e9a98"> </a>on both CPU and GPU resources vs CPU-only <a href="https://github.com/biocore/unifrac/releases/tag/0.10.0">version 0.10.0.</a></p>
SAT Competition 2004 Benchmarks and Raw Results
<p>Set of problems used for the SATISFIABILITY (SAT) competition organized in 2004. Problems are given in three categories: Industrial, Handmade (Now called "Crafted") and Random.</p>
Kiez Benchmarking Knowledge Graph Embeddings
<p>This upload contains pre-calculated Knowledge Graph Embeddings produced by our study "<a href="https://dbs.uni-leipzig.de/file/KIEZ_KEOD_2021_Obraczka_Rahm.pdf">An Evaluation of Hubness Reduction Methods for Entity Alignment with Knowledge Graph Embeddings</a>"</p>
Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity"
<p>Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity".</p> <p>Data and further information at GitHub repository https://github.com/kreutz-lab/dia-benchmarking (DOI: 10.5281/zenodo.6371925)</p>
Dataset of UAV thermal video sequences with annotations for MOTS benchmarking
<p>Instance segmentation dataset created for the research 'Monitoring Mammalian Herbivores via Convolutional Neural Networks implemented on Thermal UAV imagery'. It comprises 959 frames, 20.647 masks, and 239 tracks, and consists of 7 video sequences depicting aerial thermal imagery of cattle collected with a UAV (Parrot ANAFI Thermal) in two outdoor farms in the Netherlands. Data were acquired at three temperatures (10ºC, 19ºC, and 26.5ºC), under sunny and overcast weather conditions, at various angles of inclination (including nadir), and at heights ranging between 8-28 meters. Ground truth was labeled manually with the Computer Vision Annotation Tool <em>CVAT</em>.</p>
Input files required by the bioconvert benchmark
<p>These files required to launch the bioconvert benchmarking snakemake framework. For details see https://bioconvert.readthedocs.io .</p>
A new remote sensing benchmark dataset for machine learning applications : MultiSenGE
<p>[UPDATE] You can now access MultiSen (GE and NA) collection though this portal : <a href="https://doi.theia.data-terra.org/ai4lcc/?lang=en">https://doi.theia.data-terra.org/ai4lcc/?lang=en</a></p> <p>MultiSenGE is a new large-scale multimodal and multitemporal benchmark dataset covering one of the biggest administrative region located in the Eastern part of France. It contains 8,157 patches of 256 * 256 pixels for Sentinel-2 L2A, Sentinel-1 GRD and a regional LULC topographic regional database. </p> <p>Every file has a specific nomenclature :</p> <ul> <li>Sentinel-1 patches: {tile}_{date}_S1_{x-pixel-coordinate}_{y-pixel-coordinate}.tif</li> <li>Sentinel-2 patches: {tile}_{date}_S2_{x-pixel-coordinate}_{y-pixel-coordinate}.tif</li> <li>Ground reference patches: {tile}_GR_{x-pixel-coordinate}_{y-pixel-coordinate}.tif</li> <li>JSON Labels: {tile}_{x-pixel-coordinate}_{y-pixel-coordinate}.json</li> </ul> <p>where <em>tile</em> is the Sentinel-2 tile number, <em>date</em> the date of acquisition of the patch, <em>x-pixel-coordinate</em> and <em>y-pixel-coordinate</em> are the coordinates of the patch in the tile.</p> <p>In addition, you can find a set of useful python tools for extracting information about the dataset on Github : <a href="https://github.com/r-wenger/MultiSenGE-Tools">https://github.com/r-wenger/MultiSenGE-Tools</a></p> <p>First experiments based on this <em>dataset</em> is in press in ISPRS Annals : <strong>Wenger, R., </strong>Puissant, A., Weber, J., Idoumghar, L., and Forestier, G.: MULTISENGE: A MULTIMODAL AND MULTITEMPORAL BENCHMARK DATASET FOR LAND USE/LAND COVER REMOTE SENSING APPLICATIONS, ISPRS Ann. Photogramm. Remote Sens. Spatial Inf. Sci., V-3-2022, 635–640, https://doi.org/10.5194/isprs-annals-V-3-2022-635-2022, 2022.</p> <p>Due to the large size of the dataset, you will only find the associated JSON files on this Zenodo repository. To download the Sentinel-1, Sentinel-2 patches and the reference data, please do so via these links: </p> <ul> <li>Sentinel-1 temporal serie patches: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/s1.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/s1.tgz</a></li> <li>Sentinel-2 temporal serie patches: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/s2.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/s2.tgz</a></li> <li>Ground reference patches: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/ground_reference.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/ground_reference.tgz</a></li> <li>JSON files for each patch: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/labels.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/labels.tgz</a></li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.