Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

369

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

369 results for “Datasets, Benchmarking”

Learn how ShareScore rates datasets ↗
zenodo32/100

A benchmark dataset of diurnal- and seasonal-scale radiation, heat and CO2 fluxes in a typical East Asian monsoon region

<p>A benchmark dataset include&nbsp;30-min meteorology and&nbsp;eddy flux variables at four sites with two typical surface types (i.e., SX-cropland, DT-cropland, XZ-suburb, and DS-suburb) in the Yangtze River Delta of China.<br> SX-cropland:&nbsp;15 Jul 2015&ndash;24 Apr 2019<br> DT-cropland:&nbsp;1 Dec 2014&ndash;30 Nov 2017<br> XZ-suburb:&nbsp;27 Mar 2014&ndash;22 Jan 2017<br> DS-suburb:&nbsp;16 Apr 2011&ndash;1 Jan 2019</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Benchmark Datasets for Inductive Entity Alignment

<p>This upload contains datasets for benchmarking inductive entity alignment approaches.</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Example dataset to run end-to-end pipeline (Xenium benchmarking)

<p>This repository contains an. This example dataset is a reduced form of the one shared by Kukanja, Langseth et al. 2024. (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.cell.2024.02.030" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.cell.2024.02.030</a>) and represents a human section of spinal chord (inactive lession of Multiple schlerosis) profiled with Xenium. Please refer to the original publication for further information.</p> <p>This dataset is expected to be used as an input for the end-to-end pipeline built in the context of Marco Salas et al. 2024 (Xenium benchkmarking).&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Xenium benchmarking- xenium formatted to AnnData 2 (human breast datasets)

<p>This repository contains Xenium datasets formated to AnnData format for benchmarking the characteristics and performance of Xenium (Marco Salas et al., 2024). Please visit https://github.com/Moldia/Xenium_benchmarking to obtain further information about the datasets. On summary, AnnData contains the expression of profiled cells with some metadata, including spatial position. In adata.obs['spots'] we also include the position of decoded reads as well as some additional metadata and quality.&nbsp;</p> <p>&nbsp;</p> <p>This repository is the part 2/4 of all repositories containing the datasets used in the study already formated to AnnData</p>

opencc-by-4.0May 2024View details →
zenodo32/100

ccRCC reference datasets to benchmark UnitedMet

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

Benchmark datasets for biomedical knowledge graphs with negative statements

<p>We present a collection of datasets for three relation prediction tasks - protein-protein interaction prediction, gene-disease association prediction and disease prediction - that aim at circumventing the difficulties in building benchmarks for knowledge graphs with negative statements. These datasets include data from two successful biomedical ontologies, Gene Ontology and Human Phenotype Ontology, enriched with negative statements.&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

AlleNoise - large-scale text classification benchmark dataset with real-world label noise

<div> <div> <div> <div> <p><span>AlleNoise</span><span> is a benchmark dataset for large-scale multi-class text classification with real-world label noise. It consists of e-commerce product titles from Allegro.com with corresponding category labels. The noise distribution comes from actual users of a major e-commerce marketplace, so it realistically reflects the semantics of human mistakes. In addition to the noisy labels, we provide human-verified clean labels and a meaningful, hierarchical taxonomy of categories. Code and data is available at https://github.com/allegro/AlleNoise.<br></span></p> </div> </div> </div> </div>

opencc-by-nc-nd-4.0Jun 2024View details →
zenodo32/100

Four psychometrically validated datasets for benchmarking large language models, based on the TIMSS 2008 and 2011 released items.

<p>Four datasets validated according to psychometric principles that can be used to benchmark large language models in terms of achievements in advanced school math, advanced school physics, 8th grade math and 8th grade science.</p> <p>These four datasets are derived from items released by Trends in International Mathematics and Science Study Advanced 2008 and Trends in International Mathematics and Science Study 2011. See <a href="https://nces.ed.gov/timss/released-questions.asp">link</a>.</p> <p>For more information, see our paper <a href="https://arxiv.org/abs/2404.01799">PATCH! Psychometrics-AssisTed benCHmarking of Large Language Models: A Case Study of Mathematics Proficiency</a>.</p>

opencc-by-nc-4.0Jun 2024View details →
zenodo32/100

Datasets for benchmarking GeoFlood performance with realistic problems

<p>Datasets and configurations used in benchmarking the CPU/GPU hybrid version of GeoFlood.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

A Benchmark Dataset for Manipuri Meetei-Mayek Handwritten Character Recognition

<p>A benchmark dataset is always required for any classification or recognition system. To the best of our knowledge, no benchmark dataset exists for handwritten character recognition of Manipuri Meetei-Mayek script in <em><strong>public domain</strong></em> so far.&nbsp;Manipuri, also referred to as Meeteilon or sometimes Meiteilon, is a Sino-Tibetan language and also one of the Eight Scheduled languages of Indian Constitution. It is the official language and lingua franca of the southeastern Himalayan state of Manipur, in northeastern India. This language is also used by a significant number of people as their communicating language over the north-east India, and some parts of Bangladesh and Myanmar. It is the most widely spoken language in Northeast India after Bengali and Assamese languages.&nbsp;In this work, we introduce a handwritten Manipuri Meetei-Mayek character dataset which consists of more than 5000 data samples which were collected from a diverse population group that belongs to different age groups (from 4 years to 60 years), genders, educational backgrounds, occupations, communities from three different districts of Manipur, India (Imphal East District, Thoubal District and Kangpokpi District) during March and April 2019. Each individual was asked to write down all the Manipuri characters on one A4-size paper. The recorded responses are scanned with the help of a scanner and then each character is manually segmented from the scanned images.&nbsp;This dataset consists of segmented scanned images of handwritten Manipuri Meetei-Mayek characters (Mapi Mayek, Lonsum Mayek, Cheitap Mayek, Cheising Mayek, Khutam Mayek) of size 128X128 pixels in .JPG format as well as in .MAT format.</p>

opencc-by-4.0Dec 2018View details →
zenodo32/100

Comprehensive Structural Variant Benchmark Dataset: 1100 VCF files from long-read sequencing of 10 NCBI individuals

<p>We initially collected 10 NCBI individuals: HG002 family pedigree data (HG002 [son], HG003 [father], HG004 [mother]), the HG005 family pedigree data (HG005 [son], HG006 [father], HG007 [mother]), the NA12878 subject, the HG00096 subject, the HG00512 subject and the CHM13 subject. Then we used PacBio (CLR: Continuous Long Read, CCS: Circular Consensus Sequencing) and Nanopore (ONT) platforms, 5 aligners and 10 callers to construct the pipelines, with most parameters set to default values. After that, except for 6 invalid pipelines(pbmm2-Nanovar, lra-Picky, lra-delly, lra-NanoVar, lra-NanoSV, lra-pbsv), we obtain 1100 VCF files.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Benchmarking datasets used in the manuscript "HyLight: Strain aware assembly of low coverage metagenomes"

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo32/100

Benchmark datasets for Spotless

<p>In <strong>standards.tar.gz</strong>, the silver standards are formatted as silver_standard_<strong>i</strong>-<strong>j</strong>, where&nbsp;<strong>i&nbsp;</strong>refers to the originating single-cell dataset and j refers to the <em>synthspot</em> abundance pattern. Each directory of the silver standard contains 10 replicates. All reference datasets are in the 'reference/' folder.</p> <p><strong>Single-cell dataset:</strong></p> <ol> <li>Brain cortex</li> <li>Single-cell cerebellum</li> <li>Single-nucleus cerebellum</li> <li>Hippocampus</li> <li>Kidney</li> <li>Squamous cell carcinoma</li> <li>Melanoma</li> </ol> <p><strong>Abundance patterns:</strong></p> <ol> <li>Uniform distinct</li> <li>Diverse distinct</li> <li>Uniform overlap</li> <li>Diverse overlap</li> <li>Dominant cell type</li> <li>Partially dominant cell type</li> <li>Rare cell type</li> <li>Regionally rare cell type</li> <li>Missing cell types</li> </ol> <p>We only uploaded .rds files to keep the file size smaller. For more information on file conversion, please visit the GitHub repository.</p> <p>To use&nbsp;<strong>liver_dataset.tar.gz, </strong>please read liver_README.txt before proceeding.</p> <p><strong>dataset_visualizations.zip</strong> contain visualizations of the gold standard, silver standard, and case study datasets.</p> <p><strong>raw_results.zip </strong>contains the proportions and evaluation metrics of each method, as well as aggregated metrics used to generate plots. They can be used in tandem with the <a href="https://github.com/saeyslab/spotless-benchmark/tree/main/scripts" target="_blank" rel="noopener">scripts</a> in the GitHub repository.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Datasets of RNA structures for benchmarking entanglement detection and removal protocol using RNAspider and SPQR.

<p>Datasets of RNA structures for benchmarking entanglement detection and removal protocol using RNAspider and SPQR.</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

RNA large language models embeddings on benchmark datasets

<p>This repository contains pre-computed embeddings for several RNA sequences, using most recent Large Language Models (LLM) pre-trained on RNA sequences.&nbsp;</p> <p>compressed files for each combination of RNA-LLM models and benchmarking RNA datasets.&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

A benchmark dataset for global evapotranspiration estimation based on FLUXNET2015 from 2000 to 2022 (V1.0)

<p>Our released data mainly contains four types of data:</p> <p>(1) Half-hourly or hourly gap-filled LE data: The data are well gap-filled LE data using the novel bias-corrected RF algorithm. In the filenames, &ldquo;HH&rdquo; or &ldquo;HR&rdquo; indicate half-hourly or hourly scale data, respectively. The time information in the data files includes a pair of timestamps consistent with those in FLUXNET2015. The data are recorded at local time. The start time is &ldquo;2000-02-18, 00:00:00&rdquo;, and the end time is the same as the observation time at each site. For the quality control flags (QC), a value of 0 indicates observed data, while 1 indicates gap-filled data.</p> <p>(2) Prolonged daily LE data: This dataset provides the prolonged daily LE data using the novel bias-corrected RF algorithm. The seamless data covers the period from February 18, 2000, to December 31, 2022. For the prolonged part, the quality flag is set to 2. The rest data is consistent with the aggregated daily LE data.</p> <p>(3) Aggregated daily, monthly and yearly LE data: The hourly dataset is aggregated from the gap-filled half-hourly data to a daily scale. The start time is &ldquo;2000-02-18&rdquo;, and the end time is the same as the observation time at each site. Data quality control flags are also provided, with the values representing the percentage of hourly observations for each day. The monthly and yearly LE data are aggregated from the prolonged daily LE data. Quality control flags represent the proportion of days with more than 90% of hourly observations in a given month or a given year. No distinction is made between prolonged data and data with complete missing observations within a day. The start time for the monthly data is March 2000, and that for the yearly data is 2001.</p> <p>All files are formatted as csv files. NDVI and debiased reference variables from ERA5-Land are also provided.</p> <p>For more details of our data, please refer to a companion research article submitted to ESSD. <span>Li, W., Yao, Z., Qu, Y., Yang, H., Song, Y., Song, L., Wu, L., and Cui, Y.: A benchmark dataset for global evapotranspiration estimation based on FLUXNET2015 from 2000-2022, Earth Syst. Sci. Data, under review, 2024.</span>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Dusa benchmarking datasets

<p>The archive at <a href="https://doi.org/10.5281/zenodo.13915346">https://doi.org/10.5281/zenodo.13915346</a> contains self-contained artifacts necessary to replicate the performance evaluation; this artifact contains the repository used to generate the results in the paper, as well as the raw datasets (in the results directory) that were used to generate the charts in the papers (in the Numbers spreadsheet analysis/analysis.numbers).</p> <p>This is a repository snapshot of <a href="https://github.com/robsimmons/dusa-benchmarking/">https://github.com/robsimmons/dusa-benchmarking/</a></p>

opengpl-3.0-or-laterOct 2024View details →
zenodo32/100

Convallaria dataset for microscopy image denoising benchmark as used in Probabilistic Noise2Void paper

<p>Convallaria dataset for microscopy image denoising benchmark as used in Probabilistic Noise2Void paper (https://ieeexplore.ieee.org/document/9098336)</p>

opencc-by-4.0Feb 2020View details →
zenodo32/100

Subset of BioID dataset (https://www.bioid.com/facedb/) used for image denoising benchmark as used in DivNoising paper (https://arxiv.org/abs/2006.06072)

<p>The original BioID dataset comes from&nbsp;https://www.bioid.com/facedb/.&nbsp;</p> <p>A subset of original BioID dataset was used for image denoising benchmark (corrupted with zero mean Gaussian noise of std 15) as in DivNoising paper (https://arxiv.org/abs/2006.06072)</p>

opencc-by-4.0Jun 2020View details →
zenodo32/100

Flywing (noise 10) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)

<p>Flywing n10 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>

opencc-by-4.0May 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record