Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

369

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

369 results for “Datasets Benchmarking”

Learn how ShareScore rates datasets ↗
zenodo24/100

Mouse actin dataset for microscopy image denoising benchmark as used in PPN2V paper

<p>Mouse actin dataset for microscopy image denoising benchmark as used in PPN2V paper (https://arxiv.org/abs/1911.12291)</p>

opencc-by-4.0Nov 2019View details →
zenodo24/100

Mouse skull nuclei dataset for microscopy image denoising benchmark as used in PPN2V paper

<p>Mouse skull nuclei dataset for microscopy image denoising benchmark as used in PPN2V paper (https://arxiv.org/abs/1911.12291)</p>

opencc-by-4.0Nov 2019View details →
zenodo24/100

GNN Benchmarking Datasets

<p>Datasets appearing in &quot;https://arxiv.org/abs/2003.00982&quot; converted to HDF5</p>

opencc-by-4.0Sep 2021View details →
zenodo24/100

MIntRec2. 0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations

<p><span>MIntRec2.0, a large-scale benchmark dataset for multimodal intent recognition in multi-party conversations. It contains 1,245 high-quality dialogues with 15,040 samples, each annotated within a new intent taxonomy of 30 fine-grained classes, across text, video, and audio modalities. In addition to more than 9,300 in-scope samples, it also includes over 5,700 out-of-scope samples appearing in multi-turn contexts, which naturally occur in real-world open scenarios, enhancing its practical applicability.</span> <span>This dataset will be released under the CC BY-NC-SA 4.0 license.</span></p>

openJun 2023View details →
dryad24/100

SNP datasets and genomes used to benchmark the SNPLift program

<p>Motivation: The advent of high-throughput sequencing technologies and availability of reference genomes has provided an unprecedented opportunity to discover and genotype millions of genetic variants in hundreds or even thousands of samples. Variant calling, the identification of genetic variants from raw sequencing data, is a time-consuming and computationally expensive process. Currently, reference genomes are evolving very rapidly and new versions come out more and more frequently. To take advantage of new or improved reference genomes, raw reads alignments, genotype calling, and filtration must typically all be redone. This is a costly and time consuming operation that is not always possible when projects are under time constraints.</p> <p>Results: Here, we present SNPLift, a bioinformatic pipeline that can quickly transfer SNP coordinates from one version of a genome to another, making it possible to rapidly leverage the resources represented by new reference genomes. We tested SNPLift on nine SNP datasets in VCF format from different species (Homo sapiens, Arabidopsis thaliana, Coregonus clupeaformis, Medicato truncatula, Oriza sativa, Salvelinus namaycush, Solanum lycopersicum, Zea mays, and Glycine max). Depending on the species, we accurately lifted between 82.64% and 99.39% of the variants very quickly and efficiently, reducing the required computing power by multiple orders of magnitudes compared to a complete re-analysis using the new genome reference. SNPLift provides an accurate, parallelized, efficient and fast solution to update genome positions, for example for variant calls, based on new reference genomes.</p> <p>Availability and implementation: SNPLift is available at <a href="https://github.com/enormandeau/snplift">https://github.com/enormandeau/snplift</a> with its documentation and installation procedure. It also contains a script that runs an automated test on a small dataset, composed of 190,443 SNPs in chromosome 1 of Medicago truncatula. SNPLift uses only common tools that are easy to install and works under Linux and MacOS.</p>

opencc-zeroJun 2023View details →
ClinicalTrials.gov24/100

Age-Related Macular Degeneration Benchmark Imaging Dataset (ABID)

ClinicalTrials.gov study NCT06924021. IPD Sharing: NO. Countries: 7. Publications: 0.

closedIPD-NOFeb 2026View details →
dryad24/100

SNP datasets and genomes used to benchmark the SNPLift program

Open the record for dataset details and reuse information.

publicJun 2023View details →
geo24/100

Benchmarking of Computational Demultiplexing Methods for Single-Nucleus RNA Sequencing Data [dataset 1]

GEO Series GSE298265. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2025View details →
geo24/100

Designing a single cell RNA sequencing benchmark dataset to compare protocols and analysis methods (RNAmix_CEL-seq2 )

GEO Series GSE117617. Homo sapiens. 1 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2018View details →
geo24/100

Designing a single cell RNA sequencing benchmark dataset to compare protocols and analysis methods [5 Cell Lines 10X]

GEO Series GSE126906. Homo sapiens. 1 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2019View details →
zenodo20/100

CMFeed: A Benchmark Dataset for Controllable Multimodal Feedback Synthesis

<p><strong>Overview</strong><br>The Controllable Multimodal Feedback Synthesis (CMFeed) Dataset is designed to enable the generation of sentiment-controlled feedback from multimodal inputs, including text and images. This dataset can be used to train feedback synthesis models in both uncontrolled and sentiment-controlled manners. Serving a crucial role in advancing research, the CMFeed dataset supports the development of human-like feedback synthesis, a novel task defined by the dataset's authors. Additionally, the corresponding feedback synthesis models and benchmark results are presented in the <a href="https://github.com/MIntelligence-Group/CMFeed/" target="_blank" rel="noopener">associated code and research publication</a>.&nbsp;<br><br><em>Task Uniqueness</em>: The task of controllable multimodal feedback synthesis is unique, distinct from LLMs and tasks like VisDial, and not addressed by multi-modal LLMs. LLMs often exhibit errors and hallucinations, as evidenced by their auto-regressive and black-box nature, which can obscure the influence of different modalities on the generated responses [<a href="https://www.nature.com/articles/s41586-024-07421-0" target="_blank" rel="noopener">Ref1</a>; <a href="https://arxiv.org/abs/2311.05232" target="_blank" rel="noopener">Ref2</a>]. Our approach includes an interpretability mechanism, as detailed in the supplementary material of the corresponding <a href="https://github.com/MIntelligence-Group/CMFeed/" target="_blank" rel="noopener">research publication</a>, demonstrating how metadata and multimodal features shape responses and learn sentiments. This controllability and interpretability aim to inspire new methodologies in related fields.<br><br><strong>Data Collection and Annotation</strong><br>Data was collected by crawling Facebook posts from major news outlets, adhering to ethical and legal standards. The comments were annotated using four sentiment analysis models: FLAIR, SentimentR, RoBERTa, and DistilBERT. Facebook was chosen for dataset construction because of the following factors:<br>&bull; Facebook was chosen for data collection because it uniquely provides metadata such as news article link, post shares, post reaction, comment like, comment rank, comment reaction rank, and relevance scores, not available on other platforms.&nbsp;<br>&bull; Facebook is the most used social media platform, with 3.07 billion monthly users, compared to 550 million Twitter and 500 million Reddit users.&nbsp; [<a href="https://en.wikipedia.org/wiki/List_of_social_platforms_with_at_least_100_million_active_users" target="_blank" rel="noopener">Ref</a>] <br>&bull; Facebook is popular across all age groups (18-29, 30-49, 50-64, 65+), with at least 58% usage, compared to 6% for Twitter and 3% for Reddit. [<a href="https://sproutsocial.com/insights/new-social-media-demographics/" target="_blank" rel="noopener">Ref</a>]. Trends are similar for gender, race, ethnicity, income, education, community, and political affiliation [<a href="https://pewresearch.org/internet/fact-sheet/social-media/" target="_blank" rel="noopener">Ref</a>]&nbsp;<br>&bull; The male-to-female user ratio on Facebook is 56.3% to 43.7%; on Twitter, it's 66.72% to 23.28%; Reddit does not report this data. [<a href="https://khoros.com/resources/social-media-demographics-guide" target="_blank" rel="noopener">Ref</a>]</p> <p><em>Filtering Process</em>: To ensure high-quality and reliable data, the dataset underwent two levels of filtering:<br>a) Model Agreement Filtering: Retained only comments where at least three out of the four models agreed on the sentiment.<br>b) Probability Range Safety Margin: Comments with a sentiment probability between 0.49 and 0.51, indicating low confidence in sentiment classification, were excluded.<br>After filtering, 4,512 samples were marked as XX. Though these samples have been released for the reader's understanding, they were not used in training the feedback synthesis model proposed in the corresponding research paper.<br><br><strong>Dataset Description</strong><br>&bull; Total Samples: 61,734<br>&bull; Total Samples Annotated: 57,222 after filtering.<br>&bull; Total Posts: 3,646<br>&bull; Average Likes per Post: 65.1<br>&bull; Average Likes per Comment: 10.5<br>&bull; Average Length of News Text: 655 words<br>&bull; Average Number of Images per Post: 3.7<br><br><strong>Components of the Dataset</strong><br>The dataset comprises two main components:<br>&bull; <em>CMFeed.csv</em> File: Contains metadata, comment, and reaction details related to each post.<br>&bull; <em>Images</em> Folder: Contains folders with images corresponding to each post.<br><br><strong>Data Format and Fields of the CSV File</strong><br>The dataset is structured in CMFeed.csv file along with corresponding images in related folders. This CSV file includes the following fields:<br>&bull; <em>Id</em>: Unique identifier&nbsp;<br>&bull; <em>Post</em>: The heading of the news article.<br>&bull; <em>News_text</em>: The text of the news article.<br>&bull; <em>News_link</em>: URL link to the original news article.<br>&bull; <em>News_Images</em>: A path to the folder containing images related to the post.<br>&bull; <em>Post_shares</em>: Number of times the post has been shared.<br>&bull; <em>Post_reaction</em>: A JSON object capturing reactions (like, love, etc.) to the post and their counts.<br>&bull; <em>Comment</em>: Text of the user comment.<br>&bull; <em>Comment_like</em>: Number of likes on the comment.<br>&bull; <em>Comment_reaction_rank</em>: A JSON object detailing the type and count of reactions the comment received.<br>&bull; <em>Comment_link</em>: URL link to the original comment on Facebook.<br>&bull; <em>Comment_rank</em>: Rank of the comment based on engagement and relevance.<br>&bull; <em>Score</em>: Sentiment score computed based on the consensus of sentiment analysis models.<br>&bull; <em>Agreement</em>: Indicates the consensus level among the sentiment models, ranging from -4 (all negative) to 4 (all positive). 3 negative and 1 positive will result into -2 and 3 positives and 1 negative will result into +2.<br>&bull; <em>Sentiment_class</em>: Categorizes the sentiment of the comment into 1 (positive) or 0 (negative).<br><br><strong>More Considerations During Dataset Construction<br></strong>We thoroughly considered issues such as the choice of social media platform for data collection, bias and generalizability of the data, selection of news handles/websites, ethical protocols, privacy and potential misuse before beginning data collection. While achieving completely unbiased and fair data is unattainable, we endeavored to minimize biases and ensure as much generalizability as possible. Building on these considerations, we made the following decisions about data sources and handling to ensure the integrity and utility of the dataset:<em><br><br>&bull; Why not merge data from different social media platforms? </em>We chose not to merge data from platforms such as Reddit and Twitter with Facebook due to the lack of comprehensive metadata, clear ethical guidelines, and control mechanisms&mdash;such as who can comment and whether users' anonymity is maintained&mdash;on these platforms other than Facebook. These factors are critical for our analysis. Our focus on Facebook alone was crucial to ensure consistency in data quality and format.</p> <p><em>&bull; Choice of four news handles</em><strong>:</strong> We selected four news handles&mdash;BBC News, Sky News, Fox News, and NY Daily News&mdash;to ensure diversity and comprehensive regional coverage. These news outlets were chosen for their distinct regional focuses and editorial perspectives: BBC News is known for its global coverage with a centrist view, Sky News offers geographically targeted and politically varied content learning center/right in the UK/EU/US, Fox News is recognized for its right-leaning content in the US, and NY Daily News provides left-leaning coverage in New York. Many other news handles such as NDTV, The Hindu, Xinhua, and SCMP are also large-scale but may contain information in regional languages such as Indian and Chinese, hence, they have not been selected. This selection ensures a broad spectrum of political discourse and audience engagement.</p> <p><em>&bull; Dataset Generalizability and Bias</em><strong>:</strong> With 3.07 billion of the total 5 billion social media users, the extensive user base of Facebook, reflective of broader social media engagement patterns, ensures that the insights gained are applicable across various platforms, reducing bias and strengthening the generalizability of our findings. Additionally, the geographic and political diversity of these news sources, ranging from local (NY Daily News) to international (BBC News), and spanning political spectra from left (NY Daily News) to right (Fox News), ensures a balanced representation of global and political viewpoints in our dataset. This approach not only mitigates regional and ideological biases but also enriches the dataset with a wide array of perspectives, further solidifying the robustness and applicability of our research.</p> <p><em>&bull; Dataset size and diversity:</em> Facebook prohibits the automatic scraping of its users' personal data. In compliance with this policy, we manually scraped publicly available data. This labor-intensive process requiring around 800 hours of manual effort, limited our data volume but allowed for precise selection. We followed ethical protocols for scraping Facebook data , selecting 1000 posts from each of the four news handles to enhance diversity and reduce bias. Initially, 4000 posts were collected; after preprocessing (detailed in Section 3.1), 3646 posts remained. We then processed all associated comments, resulting in a total of 61734 comments. This manual method ensures adherence to Facebook&rsquo;s policies and the integrity of our dataset.</p> <p><strong>Ethical considerations, data privacy and misuse prevention<br></strong>The data collection adheres to Facebook&rsquo;s ethical guidelines [<a href="https://developers.facebook.com/terms/" target="_blank" rel="noopener">Ref</a>]. We manually scraped publicly available data in compliance with Facebook's ethical guidelines prohibiting automatic scraping [<a href="https://www.facebook.com/help/463983701520800" target="_blank" rel="noopener">Ref</a>]. We collected data that is publicly available, specifically corresponding to news articles that are publicly accessible following the protocols for ethically scraping Facebook data [<a href="https://webscraping.blog/how-to-scrape-facebook/" target="_blank" rel="noopener">Ref</a>]. The human-generated comments are included without identifying information. Aiming to prevent potential misuse, we have proactively designed our feedback synthesis system with an integral interpretability module. This feature is crucial as it not only helps in explaining how decisions are made within the system but also in detecting and preventing any misuse, such as the creation of misleading or manipulative content. Our original motivation for integrating this technology was to ensure that it is used responsibly and ethically, enhancing its positive impact while minimizing risks. By focusing on developing robust and transparent systems, we aim to foster trust and encourage the responsible use of technology in line with our ethical commitments.<br><br><strong>Code and Citation</strong><br>&bull; Code Repository: <a href="https://github.com/MIntelligence-Group/CMFeed/" target="_blank" rel="noopener">https://github.com/MIntelligence-Group/CMFeed/</a><br>&bull; Citing the Dataset: Users of the dataset should cite the corresponding paper described at the above GitHub Repository.<br><br><strong>License &amp; Access</strong><br>&bull; This dataset is released for academic research only and is free to researchers from educational or research institutes for non-commercial purposes.<br>&bull; Note that you are downloading this corpus at your own risk. No guarantee is provided, e.g. regarding the goodness of the corpus nor towards any subsequent effects. You may use it free of charge, and modify it as you wish, but clearly specify modifications if you pass modified material on.<br><br><strong>Contact</strong><br>Please send any questions about this dataset to:<br>&bull; Puneet Kumar (puneet.kumar@oulu.fi),<br>&bull; Sarthak Malik (sarthak_m@mt.iitr.ac.in),<br>&bull; Balasubramanian Raman (bala@cs.iitr.ac.in),<br>&bull; Xiaobai Li (xiaobai.li@zju.edu.cn).</p>

restrictedcc-by-4.0May 2024View details →
zenodo20/100

gpuPairHMM benchmark datasets

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
ClinicalTrials.gov20/100

Spine Surgery Video Observation Study. The Creation of a Benchmark Video (RGB-Depth) Dataset to Investigate the Feasibility of Developing a Markerless Tracking System for Spine Surgery.

ClinicalTrials.gov study NCT06580379. IPD Sharing: NO. Countries: 0. Publications: 0.

closedIPD-NOFeb 2026View details →
geo16/100

Benchmarking long-read RNA-sequencing technologies with LongBench: a cross-platform reference dataset profiling cancer cell lines with bulk and single-cell approaches

GEO Series GSE303762. Homo sapiens. 38 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2025View details →
zenodo16/100

Urban monthly land dynamics Sentinel-2 benchmark dataset

<p>The public data set of the paper "<strong>Time-series land cover change detection using deep learning-based temporal semantic segmentation</strong>" uses monthly synthesized Sentinel-2 for time series semantic change detection. A total of 32894 samples were collected. Each timestamp has a land cover type annotation. Anyone can use this data set to conduct further research. We will add more areas in the future.&nbsp;</p> <p><strong>Data description:</strong></p> <ol> <li>The time series length of the sample is 48, 48 months.</li> <li>Label mapping: 0 is water body, 1 is woodland, 2 is grassland, 3 is bare soil, 4 is impervious surface, 5 is cropland.</li> <li>For any implementation details, you can refer to the paper or github.</li> </ol> <p><strong>Paper citations:</strong></p> <p>He H, Yan J, Liang D, Sun Z, Li J, Wang L. Time-series land cover change detection using deep learning-based temporal semantic segmentation. Remote Sensing of Environment. 2024, 305:114101.</p>

restrictedcc-by-4.0Dec 2023View details →
zenodo16/100

FluentSigners-50: a signer independent benchmark dataset for Sign Language Processing

<p>A new large-scale Kazakh-Russian Sign Language dataset (FluentSigners-50) as a new Continuous Sign Language Recognition benchmark. &nbsp;FluentSigners-50 proposes to address three shortcomings of commonly used datasets:&nbsp;continuous signing, signer variety, and native signers. FluentSigners-50&#39;s main advantage is in its large signer variety: age (ranging from 8 to 57 years old), gender (18 male and 32 female), clothing, skin tone, body proportions, disability (deaf or hard of hearing), and fluency. Additionally, as the dataset was crowd-sourced: the participants were using a variety of their own recording devices (such as smartphones and web cameras), it resulted in a large variety of backgrounds, lighting conditions, camera quality, frame rates, camera aspect ratios, and angles. Finally, FluentSigners-50 contains recordings of 50 contributors that use sign language on a daily basis: either deaf, hard of hearing, hearing CODA (Child of Deaf Adults), and hearing SODA (Sibling of a Deaf Adult). As a result, the dataset contains a high degree of linguistic variability, including phonetic, phonological, lexical, and syntactic variations. It thus is a better training set for recognition of natural signing.</p> <p>The FluentSigners-50 dataset consists of everyday conversational phrases and sentences in KRSL, the sign language used in the Republic of Kazakhstan. KRSL is closely related to Russian Sign Language (RSL) and some other sign languages of the ex-Soviet Union. While no official research comparing KRSL with RSL exists, our observations based on our experience researching both languages are that they show a substantial lexical overlap and are entirely mutually intelligible. The sentences and phrases of FluentSigners-50 represent the following sentence types: statements, polar questions, wh-questions, and requests.</p> <p>All FluentSigners-50 contributors use sign language on a daily basis as they are either deaf (N=32), hard of hearing (N=6), hearing SODA (N=3), or hearing CODA (N=9). Native signers are signers who have been exposed to signed languages since birth because their parents are deaf. While the early acquisition may be necessary for the development of native language abilities, other factors, particularly the quality of language input, may play a role. According to this distinction, FluentSigners-50 has 30 CODA contributors (including nine hearing signers) and 20 who are not CODA (16 deaf, one hard of hearing, and three hearing SODA). Nevertheless, we decided to name our dataset FluentSigners-50 because all of our contributors use sign language daily, and it is their primary language of communication. They all came from various regions of Kazakhstan and are of different age and gender groups.</p>

restrictedFeb 2022View details →
zenodo16/100

Zagreb Calibration Benchmark Dataset

<p>Zagreb Calibration Benchmark Dataset is multi-modal high-resolution dataset geared at evaluating calibration solutions.</p>

restrictedJun 2022View details →
zenodo16/100

HiCervix: An Extensive Hierarchical Dataset and Benchmark for Cervical Cytology Classification

<p>In this paper, we release the largest three-level hierarchical cervical dataset (HiCervix), and propose a hierarchical vision transformer-based classification benchmark method (HierSwin).</p> <div> <h3>HiCervix Dataset:</h3> </div> <p>HiCervix includes 40,229 cervical cells and is categorized into 29 annotated classes. These classes are organized within a three-level hierarchical tree to capture fine-grained subtype information.</p> <p>&nbsp;</p> <div> <h3>Citation</h3> </div> <p>Please use below to cite this paper if you find our work useful in your research.</p> <p>Cai D, Chen J, Zhao J, Xue Y, Yang S, Yuan W, Feng M, Weng H, Liu S, Peng Y, Zhu J, Wang K, Jackson C, Tang H, Huang J, Wang X. HiCervix: An Extensive Hierarchical Dataset and Benchmark for Cervical Cytology Classification. IEEE Trans Med Imaging. 2024 Jun 26;PP. doi: 10.1109/TMI.2024.3419697. Epub ahead of print. PMID: 38923481.</p>

restrictedcc-by-4.0Apr 2024View details →
zenodo16/100

ArabicSL-Bench: A Benchmark Image Dataset for Arabic Alphabets Sign Language

<p>The dataset contain a total of 28,000 RGB images belonging to a total of&nbsp;28 classes. The data was collected from&nbsp;around 50 participants from Zagazig university. The ages of participants are ranging from 15 to 35 years. The images was captured by two mobile phones namely Realme 6, Realme 7,&nbsp;Realme 8. The images were resized into size of 224*224.</p> <p>The inital version of data is available privately on our page: https://www.kaggle.com/datasets/deepologylab/esl-net</p>

restrictedDec 2022View details →
zenodo16/100

IEA PVPS Task 16 Satellite nowcasting benchmark : Dataset of animated satellite images for the selection of case studies

<p>This dataset contains animated satellite images used for the selection of case studies in the framework of the IEA PVPS task 16 satellite nowcasting benchmark. The animated images contains images from the channel 12 of MSG as well as NWC SAF products (cloud types and cloud top height) over France over the year 2020.</p>

restrictedMay 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record