Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
369
datasets available to search
ShareScore release 0.9.0
Dataset results
369 results for “Datasets Benchmarking”
Supplementary dataset for "Evaluating Graph Neural Networks for Link Prediction: Current Pitfalls and New Benchmarking"
<p>The supplementary dataset for the paper "Evaluating Graph Neural Networks for Link Prediction: Current Pitfalls and New Benchmarking". We include the splits for cora, citeseer, and pubmed, the hard negative samples, and the node2vec embeddings. We also include a jupyter file <em>read_data.ipynb</em> to show how to read the non-txt file.</p> <ul> <li>heart_test_samples.npy, heart_valid_samples.npy: the heard negative samples</li> <li>*-n2v-embedding.pt: node2vec embeddings</li> <li>test_samples_index.pt, valid_samples_index.pt: the node index of the selected samples in ogbl-ppa under HeaRT</li> <li>gnn_feature: the input feature of cora, citeseer, pubmed</li> </ul> <p>More details for our code and how to use the dataset are on the code repository: https://github.com/Juanhui28/HeaRT .</p>
Datasets for Paper "BenchTemp: A General Benchmark for Evaluating Temporal Graph Neural Networks"
<p>Datasets for Paper "BenchTemp: A General Benchmark for Evaluating Temporal Graph Neural Networks"<br> URL: https://github.com/qianghuangwhu/benchtemp</p> <p>Openreview: https://openreview.net/forum?id=rnZm2vQq31</p> <p><br> There are 19 (15+4) benchmark temporal graph datasets:<br> reddit,<br> wikipedia,<br> mooc,<br> lastfm,<br> enron,<br> SocialEvo,<br> uci,<br> CollegeMsg,<br> TaobaoSmall,<br> CanParl,<br> Contacts,<br> Flights,<br> UNtrade,<br> USLegis,<br> UNvote,</p> <p>DGraphFin,</p> <p>TaobaoLarge,</p> <p>YoutubeReddit,</p> <p>YoutubeRedditLarge</p> <p> </p> <p><br> Each dataset has three files:<br> 1. ml_{data_name}.csv - the csv file of the Temporal Graph.</p> <p>This file have five columns with properties:</p> <p>'u': the id of the user.<br> 'i': the id of the item.<br> 'ts': the timestamp of the interaction (edge) between the user and the item.<br> 'label': the label of the interaction (edge).<br> 'idx': the index of the interaction (edge).<br> For example:</p> <p>,u,i,ts,label,idx<br> 0,1,2,0.0,0.0,1<br> 1,1,3,0.0,0.0,2<br> 2,1,4,0.0,0.0,3<br> 2. ml_{data_name}.npy - the edge features corresponding to the interactions (edges) in the the Temporal Graph..</p> <p>3. ml_{data_name}_node.npy - the initialization node features of the Temporal Graph.</p>
VIO-GNSS Dataset: Benchmarking Dataset for Sensor Fusion of Visual Inertial Odometry and GNSS Positioning
<p>This upload contains datasets for benchmarking and improving different Sensor Fusion implementations/algorithms. The documentation for these datasets can be found on <a href="https://github.com/AaltoVision/vio-gnss-dataset">GitHub</a>.</p> <p>The upload contains two datasets (version 1.0.0):</p> <ul> <li>urban_with_gnss_dead_zones (7.0 GB, ~16 minutes) <ul> <li>City streets</li> <li>A building is passed through on two occasions which makes the GNSS location signal unavailable at times.</li> <li>RTK Fix is acquired at times</li> </ul> </li> <li>suburban_nature (10.6 GB, ~19 minutes) <ul> <li>The route begins on a suburban street but quickly turns into a nature trail. Lots of vegetation</li> <li>The RTK solution is only Float or None most of the route.</li> </ul> </li> </ul> <p>Details on collecting the data:</p> <ul> <li>Software <ul> <li>The data was collected using <a href="https://github.com/AaltoVision/vio-gnss-recorder">this</a> open-source recorder. <ul> <li>Can be easily replayed using <a href="https://github.com/SpectacularAI/sdk-examples">SpectacularAI's SDK</a> (sdk-examples/python/oak/vio_replay.py)</li> </ul> </li> <li>Each dataset contains a map of the travelled route in Otaniemi, Espoo, Finland.</li> <li><strong>Necessary files to implement SLAM are included</strong> in the dataset.</li> <li>Use of NTRIP and the high precision GNSS antenna enables global positioning accuracy of only few centimeters.</li> </ul> </li> <li>Hardware <ul> <li>OAK-D stereo depth + color camera (Luxonis)</li> <li>C099-F9P GNSS module (u-blox)</li> <li>ANN-MB-00 high precision GNSS antenna (u-blox)</li> </ul> </li> </ul>
Benchmarking of PROTAC docking and virtual screening tools - dataset
<p>This repository includes all the input files for the PROTAC docking and virtual screening benchmark. The raw data files are available on request (all raw data files combined is ~70GB and compressed ~47GB). Researchers can access and download the data to reproduce the results. Details about files/folders is included in the README.txt.</p>
FOR-instance: a UAV laser scanning benchmark dataset for semantic and instance segmentation of individual trees
<p>The challenge of accurately segmenting individual trees from laser scanning data hinders the assessment of crucial tree parameters necessary for effective forest management, impacting many downstream applications. While dense laser scanning offers detailed 3D representations, automating the segmentation of trees and their structures from point clouds remains difficult. The lack of suitable benchmark datasets and reliance on small datasets have limited method development. The emergence of deep learning models exacerbates the need for standardized benchmarks. Addressing these gaps, the FOR-instance data represent a novel benchmarking dataset to enhance forest measurement using dense airborne laser scanning data, aiding researchers in advancing segmentation methods for forested 3D scenes.</p> <p>In this repository, users will find forest laser scanning point clouds from unamnned aerial vehicle (using Riegl sensors) that are manually segmented according to the individual trees (1130 trees) and semantic classes. The point clouds are subdivided into five data collections representing different forests in Norway, the Czech Republic, Austria, New Zealand, and Australia. </p> <p>These data are meant to be used either for developement of new methods (using the dev data) or for testing of exisitng methods (test data). The data splits are provided in the data_split_metadata.csv file.</p> <p>A full description of the FOR-instance data can be found at <a href="http://arxiv.org/abs/2309.01279">http://arxiv.org/abs/2309.01279</a> </p>
Small version of other JSON benchmarking datasets
<p>10% size version of bestbuy, google, twitter, and walmart datasets used for benchmarking JSON engines.<br> <br> Full versions:<br> <br> https://zenodo.org/record/7607865</p> <p>https://zenodo.org/record/7607889</p> <p>https://zenodo.org/record/7607891</p> <p>https://zenodo.org/record/7607882</p>
Evaluation results of the xMEN entity linking toolkit for multiple benchmark datasets
Open the record for dataset details and reuse information.
NA12878 WES Benchmark dataset
<p>This dataset makes available the UCSC Genome Browser (genome.ucsc.edu) GRCh37 genome build public session <strong>NA12878 WES Benchmark </strong>files in a single dataset so that these files can be used in other applications or genome browsers such as IGV. </p> <p>The <a href="https://usegalaxy.org/u/erinija/p/omim-genes-in-na12878-wes-benchmark">"Procedure and datasets to cross-reference OMIM genes with the genomic regions of interest"</a> Galaxy page on <strong>usegalaxy.org</strong> server's <em><strong>Shared Data Pages</strong></em> describes practical procedure and several possible use cases for this data set. This page can be accessed freely by users logged into their accounts on usegalaxy.org. Please register if you don't have an account on usegalaxy.org Galaxy server. </p> <p>All genomic variant calls in all VCF files of this data set were decomposed and normalized with vt. This dataset contains: </p> <ol> <li>Genome in a bottle (GIAB) version 3.3.2 high confidence (HC) variant calls and genomic regions for HapMap individual NA12878 : <ol> <li>GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz</li> <li>GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz.tbi</li> <li>GIAB_v3.3.2_NA12878_HC_regions.bed</li> </ol> </li> <li>HapMap individual NA12878 WES variant calls (VCF) and capture regions (BED) from diagnostic laboratories : <ul> <li>ARUP whole exome sequencing data (HiSeq 2000) publically available from NCBI GeT-RM Browser <ol> <li>converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz</li> <li>converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz.tbi</li> <li> ARUP_SeqCap_EZ_Exome.bed</li> </ol> </li> <li>UCSF whole exome sequencing data (HiSeq 2500) publically available from NCBI GeT-RM Browser <ol> <li>converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz</li> <li>converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz.tbi</li> <li>UCSF_WES_Agilent_V4_Custom.bed</li> </ol> </li> <li>Whole exome data (NextSeq 500) sequenced in CHEO diagnostic laboratory <ol> <li>CHEO_NA12878_WES_S1dataset.vcf.gz</li> <li>CHEO_NA12878_WES_S1dataset.vcf.gz.tbi</li> <li>Agilent_CRE_v2.bed</li> </ol> </li> </ul> </li> <li>Genomic coordinates (BED) of OMIM genes for which a molecular basis of the associated disease is known (as of September 2019) : <ul> <li>Omim_Genes.bed </li> </ul> </li> </ol>
ClaimBuster: A Benchmark Dataset of Check-worthy Factual Claims
<p>The ClaimBuster dataset consists of statements extracted from all U.S. general election presidential debates (1960-2016) along with human-annotated check-worthiness labels where each sentence is categorized into one of the three categories: non-factual statement, unimportant factual statement, and check-worthy factual statement.</p> <p>If you use this dataset, please cite the following paper:</p> <p>@inproceedings{arslan2020claimbuster,<br> title={{A Benchmark Dataset of Check-worthy Factual Claims}},<br> author={Arslan, Fatma and Hassan, Naeemul and Li, Chengkai and Tremayne, Mark },<br> booktitle={14th International AAAI Conference on Web and Social Media},<br> year={2020},<br> organization={AAAI}<br> }</p>
RibFrac Dataset: A Benchmark for Rib Fracture Detection, Segmentation and Classification (Test Set)
<p>RibFrac dataset is a benchmark for developping algorithms on rib fracture detection, segmentation and classification. We hope this large-scale dataset could facilitate both clinical research for automatic rib fracture detection and diagnoses, and engineering research for 3D detection, segmentation and classification.</p> <p>This is the Test Set of RibFrac dataset, including 160 CTs and the corresponding annotations. Files include:</p> <ol> <li>ribfrac-test-images.zip: 160 chest-abdomen CTs in NII format (nii.gz) </li> </ol> <p>Note that only the images are released; the corresponding ground truth labels are not available. Check the <a href="https://ribfrac.grand-challenge.org/">MICCAI 2020 RibFrac Challenge website</a> to evaluate your algorithm.</p> <p> </p> <p>If you find this work useful in your research, please acknowledge the RibFrac project teams in the paper and cite this project as:</p> <p><em>Liang Jin, Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yiyi Gao, Yingli Sun, Pan Gao, Weiling Ma, Mingyu Tan, Hui Kang, Jiajun Chen, Ming Li. Deep-</em><em>Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet. EBioMedicine (2020). (<a href="https://doi.org/10.1016/j.ebiom.2020.103106">DOI</a>)</em></p> <p>or using bibtex</p> <p><em>@article{ribfrac2020,<br> title={Deep-Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet},<br> author={Jin, Liang and Yang, Jiancheng and Kuang, Kaiming and Ni, Bingbing and Gao, Yiyi and Sun, Yingli and Gao, Pan and Ma, Weiling and Tan, Mingyu and Kang, Hui and Chen, Jiajun and Li, Ming},<br> journal={EBioMedicine},<br> year={2020},<br> publisher={Elsevier}<br> }</em></p> <p> </p> <p>The RibFrac dataset is a research effort of thousands of hours by experienced radiologists, computer scientists and engineers. We kindly ask you to respect our effort by appropriate citation and keeping data license.</p> <p> </p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>
WCET benchmarks test on STM32, Raspberry PI/Linux and Raspberry PI/PREEMPT_RT dataset
<p>Please check the README file.</p>
Multiplane microscopy dataset for benchmarking denoising methods
<p>The dataset consists of 810 microscopy images collected from CHO, U2OS, and RPE1 cell lines in fluorescent and brightfield modalities. Cell lines were grown in standard medium, seeded on plates, fixed with formaldehyde, stained with a fluorescent dye (Hoechst33342) and imaged with PerkinElmer Phenix high-throughput confocal microscope in fluorescence and brightfield modalities using a 20x objective. For the fluorescent modality, images were acquired in low-exposure (20 ms) and high-exposure (100 ms) modes. For the brightfield, we used a single exposure (100 ms). Higher exposure time usually translates into better image quality, alleviating Poisson noise, however, it leads to sample degradation, and lower measurement speed.</p> <p>In total, three wells on a plate were imaged, one well per cell line, each with nine fields of view. Each field of view was imaged across ten focal planes separated by 2 micrometers in z-stack, resulting in 270 microscopy images of size 2160 x 2160 pixels per exposure time per modality.</p> <p>We used the following folder structure: /{well}_{cell line}/{modality}_{exposure time}/fov_{field of view}_plane_{plane}.png</p> <p>We recommend to use the first two wells (CHO and U2OS, 180 images) as a training set, the first three fields of view from the last well (RPE1, 30 images) as a validation set, and the last six fields of view (RPE1, 60 images) as a test set.</p>
Heteroplasmy Benchmark Dataset - mitochondrial DNA mixture model - MiSeq - U5-H1-M1-M2-M3-M4-M5 - BAM
<p>mtDNA mixture model of 2 mtDNA sequences belonging to haplogroups U5 and H1. Run on Illumina MiSeq with 3 different polymerases (Clontech, Herculase, NEB Taq), and different DNA extraction protocols - <strong>BAM FILES </strong></p> <p>M1 = Mixture 1:2 i.e. 50%</p> <p>M2 = Mixture 1:10 i.e. 10%</p> <p>M3 = Mixture 1:50 i.e. 2%</p> <p>M4 = Mixture 1:100 i.e. 1%</p> <p>M5 = Mixture 1:200 i.e. 0.5%</p>
Benchmark dataset of MULTITEST - TEMP12
<p>As a part of the Spanish MULTITEST project (Multiple verification of automatic softwares homogenizing monthly temperature and precipitation series, <strong>CGL2014-52901-P</strong>, 2015-2017) 12 monthly temperature benchmark datasets were examined with 9 versions of 5 homogenization methods. Here we publish the data of these experiments in text files.</p> <p>See more details in the attached "Readme".</p> <p> </p> <p> </p>
AMDADOS datasets and scaling benchmarks
<p>This dataset provides data from benchmark scaling experiments on the AMDADOS advection diffusion code integrated with data assimilation. Results are included for performance in serial environment, MPI parallel and parallelized using the AllScale environment</p>
CoqStoq: A Dataset of Coq Proofs Scraped from GitHub and a Corresponding Benchmark
<p>CoqStoq is a dataset of Coq Proofs scraped from GitHub. CoqStoq contains proofs from 2,205 open source projects and includes a benchmark on proofs from 12 projects. This repository includes the projects from CoqStoq's benchmarks and the processed proof data collected by CoqStoq.</p>
Re-assembling the past: The RePAIR dataset and benchmark for real world 2D and 3D puzzle solving
<p><em>Accepted by NeurIPS 2024 Datasets and Benchmarks Track</em></p> <p>We introduce the RePair puzzle-solving dataset, a large-scale real world dataset of fractured frescoes from the archaelogical campus of Pompeii. Our dataset consists of over 1000 fractured frescoes. The RePAIR stands as a realistic computational challenge for methods for 2D and 3D puzzle solving, and serves as a benchmark that enables the study of fractured object reassembly and presents new challenges for geometric shape understanding. Please visit <a href="https://repairproject.github.io/RePAIR_dataset/">our website</a> for more dataset information, access to source code scripts and for an interactive gallery viewing of the dataset samples.</p> <div> <h3>Access the entire dataset</h3> <p>We provide a compressed version of our dataset in two seperate files. One for the 2D version and one for the 3D version.</p> <p>Our full dataset contains over one thousand individual fractured fragments divided into groups with its corresponding folder and all compressed into their individual sub-set format regarding whether they are 2D or 3D. Regarding the 2D dataset, each fragment is saved as a .PNG image and each group has the corresponding ground truth transformation to solve the puzzle as a <strong><em>.TXT</em></strong> file. Considering the 3D dataset, each fragment is saved as a mesh using the widely <strong><em>.OBJ</em></strong> format with the corresponding material (<strong><em>.MTL</em></strong>) and texture (<strong><em>.PNG</em></strong>) file. The meshes are already in the assembled position and orientation, so that no additional information is needed. All additional metadata information are given as <strong><em>.JSON</em></strong> files.</p> <p> </p> <h1>Important Note</h1> <p><strong>Please be advised that downloading and reusing this dataset is permitted only upon acceptance of the following license terms.</strong></p> <p><em><strong>The Istituto Italiano di Tecnologia (IIT) declares, and the user (“User”) acknowledges, that the "RePAIR puzzle-solving dataset" contains 3D scans, texture maps, rendered images and meta-data of fresco fragments acquired at the Archaeological Site of Pompeii. IIT is authorised to publish the RePAIR puzzle-solving dataset herein only for scientific and cultural purposes and in connection with an academic publication referenced as Tsemelis et al., "Re-assembling the past: The RePAIR dataset and benchmark for real world 2D and 3D puzzle solving", NeurIPS 2024. Use of the RePAIR puzzle-solving dataset by User is limited to downloading, viewing such images; comparing these with data or content in other datasets. User is not authorised to use, in particular explicitly excluding any commercial use nor in conjunction with the promotion of a commercial enterprise and/or its product(s) or service(s), reproduce, copy, distribute the RePAIR puzzle-solving dataset. User will not use the RePAIR puzzle-solving dataset in any way prohibited by applicable laws. RePAIR puzzle-solving dataset therein is being provided to User without warranty of any kind, either expressed or implied. User will be solely responsible for their use of such RePAIR puzzle-solving dataset. In no event shall IIT be liable for any damages arising from such use.</strong></em></p> </div>
EuroSDR RPAS benchmark datasets
<p><strong>Overview</strong></p> <p>The <a href="https://www.eurosdr.net/">EuroSDR</a> RPAS benchmark datasets were aqcuired in August 2021 as part of the EuroSDR benchmark initiative. This aims to evaluate the true geometric quality of real-world survey data generated from Remotely Piloted Aircraft System (RPAS) photogrammetry and lidar under different control configurations, focussing primarily on the geometric quality of data generated in the absence of ground control and local GNSS base station information.</p> <p>Guided by a task force of National Mapping and Cadastral Agencies (NMCAs) experts and academics, in August 2021 Newcastle Geospatial Engineering team have established and surveyed a coordinated test field of independent checkpoints (CPs), test surfaces and profiles at the disused Wards Hill Quarry near Morpeth, Northumberland, UK. The 350 x 250 m study area was simultaneously surveyed using the following RPAS mounted instruments, each limited to a single flight to represent “real-world” operation:</p> <ul> <li>DJI Phantom 4 RTK </li> <li>DJI ZENMUSE P1 </li> <li>DJI ZENMUSE L1 </li> <li>Routescene LidarPod </li> <li>Riegl MiniVUX</li> </ul> <p>More information can be found <a href="https://geospatialncl.github.io/eurosdr-rpas-benchmark/">here</a>.</p> <p>To facilitate and encourage wider use of the EuroSDR RPAS dataset, it has been made open access.</p> <p><strong>About us</strong></p> <p>We are the <a href="https://www.ncl.ac.uk/engineering/research/civil-engineering/geospatial-engineering/">Geospatial Engineering</a> research group in the <a href="https://www.ncl.ac.uk/engineering/">School of Engineering</a> at <a href="https://www.ncl.ac.uk/">Newcastle University</a> with a long history of research and teaching across geospatial disciplines. <a href="https://www.eurosdr.net/">EuroSDR</a> is a not-for-profit organisation linking National Mapping and Cadastral Agencies with Research Institutes and Universities in Europe for the purpose of applied research in spatial data provision, management and delivery.</p> <p> </p>
APIBench: A Benchmark Dataset for Evaluating API Recommendation Approaches in Python and Java
<p>APIBench is the benchmark dataset APIBench released in the paper "<a href="https://yunpeng.site/files/apirec.pdf"><em>Revisiting, Benchmarking and Exploring APIRecommendation: How Far Are We?</em></a>". </p> <p>APIBench contains two sub-dataset for evaluating the performance of query-based and code-based API recommendation approaches, namely APIBench-Q and APIBench-C. Each sub-dataset has a Java version and a Python version.</p> <p>APIBench-Q contains 4,309 Python queries and 6,563 Java queries collected from Stack Overflow posts generated from Aug 2008 to Feb 2021 and tutorial websites Geeks4Geeks, Java2s, and Kode Java in April 2021.</p> <p>APIBench-C contains 2,361 Python projects and 1,477 Java projects mined from GitHub in April 2021.</p> <p>Please read the <strong>README.md</strong> file for detailed information about the benchmark.</p> <p>The evaluation results of existing API recommendation approaches can be found in <a href="https://github.com/JohnnyPeng18/APIBench">this GitHub repository</a>.</p>
Performance Measurement Dataset of the HPC Benchmarks FASTEST, Kripke, LULESH, MiniFE, Quicksilver, and RELeARN for Scalability Studies with Extra-P
<p>This dataset contains performance measurements of the HPC benchmarks FASTEST, Kripke, LULESH, MiniFE, Quicksilver, and RELeARN intended to be used for scalability studies with Extra-P (https://github.com/extra-p/extrap). The datasets contains measurements of various application configurations considering several model parameters, e.g., the number of MPI ranks and the input problem size, using weak scaling for each benchmark.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.