Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

369

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

369 results for “Datasets, Benchmarking”

Learn how ShareScore rates datasets ↗
zenodo40/100

MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Test Set)

<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset&nbsp;could facilitate the research and application of automatic mediastinal lesion detection and diagnosis.&nbsp;</p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is the Test Set&nbsp;of MELA dataset, including 220&nbsp;CTs. Files include:</p> <ol> <li>Test1.zip: 110 CTs in NII format (nii.gz).</li> <li>Test2.zip: 110 CTs in NII format (nii.gz).</li> </ol>

opencc-by-4.0Jun 2022View details →
zenodo40/100

PIGSFLI Benchmarking Dataset

<p>This repository contains all raw unprocessed quantum Monte Carlo data utilized in testing and benchmarking the <code>pigsfli</code> code available at <a href="https://github.com/DelMaestroGroup/pigsfli ">https://github.com/DelMaestroGroup/pigsfli</a>.</p> <p><strong>Directory Names</strong></p> <p>Directory names are encoded according to the following rule:</p> <pre>{dimension}D_{linear_size}_{total_particles}_{partition size}_{interactionpotential}_{tunneling parameter}_{beta}_{number of bins}</pre> <p><strong>File Names</strong></p> <ol> <li>Contains the state of the RNG: <pre>1D_16_16_8_7.071100_1.000000_12.000000_10001_rng-state_0_square_2.dat {dimension}D_{linear_size}_{total_particles}_{partition size}_{interaction potential}_{tunneling parameter}_{beta}_{number of bins}_rng-state_{seed}_{geometry of subregion}_{number_of_replicas}.dat </pre> </li> <li>Contains the state of the system: <pre>1D_16_16_8_7.071100_1.000000_12.000000_10001_system-state_0_square_2.dat {dimension}D_{linear_size}_{total_particles}_{partition size}_{interaction potential}_{tunneling parameter}_{beta}_{number of bins}_system-state_{seed}_{geometry of subregion}_{number_of_replicas}.dat</pre> </li> <li>Number of times each possible number of swapped sites was measured (each column is a number of swaps ranging from 0 to ℓ): <pre>1D_16_16_8_7.071100_1.000000_12.000000_10001_SWAP_137_square.dat {dimension}D_{linear_size}_{total_particles}_{partition size}_{interaction potential}_{tunneling parameter}_{beta}_{number of bins}_SWAP_{seed}_{geometry of subregion}.dat</pre> </li> <li>For fixed number of swapped sites (mA), how many times each possible local particle number was measured (columns range from n=0,...,N): <pre>1D_8_8_4_3.300000_1.000000_0.600000_10000_SWAPn-mA4_42_square.dat {dimension}D_{linear_size}_{total_particles}_{partition size}_{interaction potential}_{tunneling parameter}_{beta}_{number of bins}_SWAPn-mA{number of swapped sites}_{seed}_{geometry of subregion}.dat</pre> </li> <li>For fixed number of subregion sites (mA), how many times each possible local particle number was measured (columns range from n=0,...,N): <pre>1D_8_8_4_3.300000_1.000000_0.600000_10000_Pn-mA4_42_square.dat {dimension}D_{linear_size}_{total_particles}_{partition size}_{interaction potential}_{tunneling parameter}_{beta}_{number of bins}_Pn-mA{number of swapped sites}_{seed}_{geometry of subregion}.dat </pre> </li> <li>For fixed number of subregion sites (mA), how many times each possible local particle number was measured simultaneously on both replicas when there were no swapped sited (columns range from n=0,...,N): <pre> 1D_8_8_4_3.300000_1.000000_0.600000_10000_PnSquared-mA4_42_square.dat {dimension}D_{linear_size}_{total_particles}_{partition size}_{interaction potential}_{tunneling parameter}_{beta}_{number of bins}_PnSquared-mA{number of swapped sites}_{seed}_{geometry of subregion}.dat</pre> </li> <li>Kinetic energy: <pre> 1D_8_8_4_3.300000_1.000000_2.000000_10000_K_93_square.dat {dimension}D_{linear_size}_{total_particles}_{partition size}_{interaction potential}_{tunneling parameter}_{beta}_{number of bins}_K_{seed}_{geometry of subregion}.dat </pre> </li> <li>Potential energy: <pre>1D_8_8_4_3.300000_1.000000_2.000000_10000_V_93_square.dat {dimension}D_{linear_size}_{total_particles}_{partition size}_{interaction potential}_{tunneling parameter}_{beta}_{number of bins}_V_{seed}_{geometry of subregion}.dat</pre> </li> </ol>

opencc-by-4.0Jul 2022View details →
zenodo40/100

WaterBench-Iowa: A Large-scale Benchmark Dataset for Data-Driven Streamflow Forecasting

<p>WaterBench-Iowa is&nbsp;a comprehensive benchmark dataset for streamflow forecasting.&nbsp;It&nbsp;follows FAIR data principles that are prepared with a focus on convenience for utilizing in data-driven and machine learning studies and provides benchmark performance for state-of-art deep learning architectures on the dataset for comparative analysis. By aggregating the datasets of streamflow, precipitation, watershed area, slope, soil types, and evapotranspiration from federal agencies and state organizations (i.e., NASA, NOAA, USGS, and Iowa Flood Center), we provided the WaterBench for hourly streamflow forecast studies. This dataset has a high temporal and spatial resolution with rich metadata and relational information, which can be used for varieties of deep learning and machine learning research.&nbsp;To some extent, WaterBench makes up for the lack of a unified benchmark in earth science research. We highly encourage researchers to use the WaterBench for deep learning research in hydrology.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Historical time-series reconstruction benchmark dataset of Landsat bi-monthly aggregates from GLAD ARD-2 at 30-m resolution with stratified sampling based on ESA CCI

<h2>Description</h2> <p>Historical time-series reconstruction benchmark dataset presented here is designed for evaluating and comparing the performance of time series reconstruction methods in the context of land cover change detection. The dataset is based on the European Space Agency Climate Change Initiative (ESA CCI) land cover dataset, which has been aggregated into 18 classes to facilitate analysis. The dataset includes information on land cover dynamics from 2000 to 2020, focusing on identifying and characterizing changes in land cover over time.</p> <h3><strong>Data Collection and Processing:</strong></h3> <p>The dataset is derived from the ESA CCI land cover dataset, which provides information on land cover classes at a global scale. The original dataset, containing 37 land cover classes, was aggregated into 18 classes based on similarity. Pixels with stable land cover over the study period and pixels with one or multiple land cover changes were identified and grouped into strata for sampling purposes.</p> <p>Sampling points were selected using a stratified sampling design, ensuring representation across different land cover classes and change scenarios. Approximately 2600 points were selected from each stratum, resulting in a total of 51,978 sampling points. The selected points were uniformly distributed along the strata, with spatial variations accounted for.</p> <p>Bimonthly time series data were extracted for each sampling point from 1997 to 2022, capturing temporal dynamics in land cover. Artificial gaps were introduced into the time series data to simulate real-world data loss, allowing for the evaluation of time series reconstruction methods under varying gap densities.</p> <p>The time series values were extracted from Landsat GLAD imagery using the specified spectral bands, including blue, green, red, NIR, SWIR1, SWIR2, and thermal bands. Additionally, a clear quality band was also extracted.</p> <h3>Data Details</h3> <ul> <li><strong>Time Period:</strong> 1997-01-01 to 2022-12-31</li> <li><strong>Type of Data: </strong>R data frame / Geopackage points.</li> <li><strong>Collection/Derivation:</strong> Derived from Landsat ARD v2, processed with Scikit-map.</li> <li><strong>Coordinate Reference System:</strong> EPSG:4326</li> <li><strong>Bounding Box:</strong> All the globe</li> <li><strong>File Format:</strong> RDS</li> </ul> <p>&nbsp;</p> <h3><strong>Reclassified Classes of ESA CCI Land Cover Dataset</strong></h3> <table> <tbody> <tr> <td> <div> <div> <p><strong>Aggregated Class Code</strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Aggregated Class Label</strong></p> </div> </div> </td> <td> <div> <div> <p><strong>Original ESA CCI Classes</strong></p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>10</p> </div> </div> </td> <td> <div> <div> <p>Cropland rainfed</p> </div> </div> </td> <td> <div> <div> <p>10, 11, 12</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>30</p> </div> </div> </td> <td> <div> <div> <p>Mosaic cropland | natural vegetation</p> </div> </div> </td> <td> <div> <div> <p>30, 40</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>50</p> </div> </div> </td> <td> <div> <div> <p>Tree cover broadleaved evergreen</p> </div> </div> </td> <td> <div> <div> <p>50</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>60</p> </div> </div> </td> <td> <div> <div> <p>Tree cover broadleaved deciduous</p> </div> </div> </td> <td> <div> <div> <p>60, 61, 62</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>70</p> </div> </div> </td> <td> <div> <div> <p>Tree cover needleleaved evergreen</p> </div> </div> </td> <td> <div> <div> <p>70, 71, 72</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>80</p> </div> </div> </td> <td> <div> <div> <p>Tree cover needleleaved deciduous</p> </div> </div> </td> <td> <div> <div> <p>80, 81, 82</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>90</p> </div> </div> </td> <td> <div> <div> <p>Tree cover mixed leaf type</p> </div> </div> </td> <td> <div> <div> <p>90</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>100</p> </div> </div> </td> <td> <div> <div> <p>Mosaic tree and shrub | herbaceous cover</p> </div> </div> </td> <td> <div> <div> <p>100, 110</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>120</p> </div> </div> </td> <td> <div> <div> <p>Shrubland</p> </div> </div> </td> <td> <div> <div> <p>120, 121, 122</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>150</p> </div> </div> </td> <td> <div> <div> <p>Sparse vegetation</p> </div> </div> </td> <td> <div> <div> <p>150, 151, 152, 153</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>160</p> </div> </div> </td> <td> <div> <div> <p>Tree cover flooded</p> </div> </div> </td> <td> <div> <div> <p>160, 170</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>180</p> </div> </div> </td> <td> <div> <div> <p>Shrub or herbaceous cover flooded</p> </div> </div> </td> <td> <div> <div> <p>180</p> </div> </div> </td> </tr> <tr> <td> <div> <div> <p>200</p> </div> </div> </td> <td> <div> <div> <p>Bare areas</p> </div> </div> </td> <td> <div> <div> <p>200, 201, 202</p> </div> </div> </td> </tr> </tbody> </table> <p>In the table, each row represents a reclassified land cover class, identified by a unique code. The 'Original ESA CCI Classes' column lists the specific land cover classes from the European Space Agency Climate Change Initiative dataset that are grouped together to form each broader category. Note that land cover classes not listed in this table were retained in their original value and were not reclassified.</p> <h3><strong>File Format</strong></h3> <p>The dataset comprises observations spanning from January 1997 to November 2022, capturing data for 51,978 samples.</p> <ul> <li>blue.rds: Time series data for the blue spectral band.</li> <li>green.rds: Time series data for the green spectral band.</li> <li>red.rds: Time series data for the red spectral band.</li> <li>nir.rds: Time series data for the near-infrared (NIR) spectral band.</li> <li>swir1.rds: Time series data for the shortwave infrared 1 (SWIR1) spectral band.</li> <li>swir2.rds: Time series data for the shortwave infrared 2 (SWIR2) spectral band.</li> <li>thermal.rds: Time series data for the thermal infrared band.</li> <li>clear.rds: Time series data for the clear quality band, used for masking out cloudy observations.</li> </ul> <p>How open the files in R:</p> <p><code>blue &lt;- readRDS("blue.rds")</code></p> <p>To open the files in Python, you need to the <code>pyreadr</code> library:</p> <p><code>import pyreadr</code><br><code>blue = pyreadr.read_r('blue.rds')</code></p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Dataset for 'The State of Data Curation at NeurIPS: An Assessment of Dataset Development Practices in the Datasets and Benchmarks Track'

<p>This packages comprises of analyses and evaluations of 60 datasets from the NeurIPS Datasets and Benchmarks track. It is part of a paper currently under review at the 2024 the NeurIPS Datasets and Benchmarks track, titled, "The State of Data Curation at NeurIPS: An Assessment of Dataset Development Practices in the Datasets and Benchmarks Track".&nbsp;</p>

opencc-by-sa-4.0Jun 2024View details →
zenodo40/100

Benchmark datasets for RDF load time evaluation (RiverBench)

<div> <p>Datasets to be used for reproducing the RDF load time benchmark, using the code here: <a href="https://github.com/Ostrzyciel/rdf4led-riverbench">https://github.com/Ostrzyciel/rdf4led-riverbench</a></p> <p>The datasets were obtained from <a href="https://w3id.org/riverbench/v/2.0.1/profiles/flat-triples" rel="nofollow">RiverBench profile <code>flat-triples</code> version 2.0.1</a>. <strong>The detailed licensing and authorship information for each individual dataset is available on <a href="https://w3id.org/riverbench/v/2.0.1/datasets" rel="nofollow">RiverBench's website</a>.</strong> The most restrictive license that applies to any of the datasets is CC BY-SA.</p> <p>Benchmark results: <a href="https://doi.org/10.5281/zenodo.12087112">https://doi.org/10.5281/zenodo.12087112</a></p> </div>

opencc-by-sa-4.0Jun 2024View details →
zenodo40/100

Scorpio Gene-Taxa Benchmark Dataset

<div> <div> <div> <div>&nbsp;</div> </div> </div> </div> <div> <div> <div> <div> <div> <div> <p>We used the Woltka pipeline to compile the complete Basic genome dataset, consisting of 4634 genomes, with each genus represented by a single genome. After downloading all coding sequences (CDS) from the NCBI database, we extracted 8 million distinct CDS, focusing on bacteria and archaea and excluding viruses and fungi due to inadequate gene information.</p> <p>To maintain accuracy, we excluded hypothetical proteins, uncharacterized proteins, and sequences without gene labels. We addressed issues with gene name inconsistencies in NCBI by keeping only genes with more than 1000 samples and ensuring each phylum had at least 350 sequences. This resulted in a curated dataset of 800,318 gene sequences from 497 genes across 2046 genera.</p> <p>We created four datasets to evaluate our model: a training set (Train_set), a test set (Test_set) with different samples but the same genus and gene as the training set, a Taxa_out_set excluding 18 phyla present in the training set but from different phyla, and a Gene_out_set excluding 60 genes from the training set but from the same phyla. We ensured each CDS had only one representation per genome, removing genes with multiple representations within the same species.</p> </div> </div> </div> </div> </div> </div>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Datasets for "Precision and accuracy of single-molecule FRET measurements – a multi-laboratory benchmark study"

<p>Supplementary material (raw data) for Fig. 2 in &quot;<strong>Precision and accuracy of single-molecule FRET measurements &ndash; a multi-laboratory benchmark study</strong>&quot; to be published with Nature Methods</p> <p>The confocal data is given in ht3 and hdf5 format.</p> <p>For the TIRF data the original TIFF-stacks are uploaded including the calibration files.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

BenchPS: A Benchmark Dataset for Phrase Simplification

<p>BenchPS is a dataset built for the training and evaluation of phrase simplification systems. Each instance is composed of a sentence, target complex phrase, and a set of candidate simplifications ranked by simplicity. Each instance was annotated by humans through multiple annotations steps to ensure the reliability of the data.</p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

A Multi-Challenge Clustering Benchmark Dataset Embedding Large Differences in Spatial Extent

<p>This artificial clustering benchmark dataset was designed manually and draws its inspiration from structural aspects that can be seen in principal component plots of hyperspectral image data. Distance-separated, density-separated, gradient-separated as well as connected clusters have been placed into the dataset. Following the notion that clusters may vary significantly with respect to their spatial extent the respective separability problems are scaled at different levels and only become visible by magnifying certain parts of the dataset. Another special aspect of this dataset is that cluster borders have been kept rather ambiguous which, in our opinion, better resembles the situation in spectroscopic data.</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems

<p>Music recommender systems can offer users personalized and contextualized recommendation and are therefore important for music information retrieval. An increasing number of datasets have been compiled to facilitate research on different topics, such as content-based, context-based or next-song recommendation. However, these topics are usually addressed separately using different datasets, due to the lack of a unified dataset that contains a large variety of feature types such as item features, user contexts, and timestamps. To address this issue, we propose a large-scale benchmark dataset called #nowplaying-RS, which contains 11.6 million music listening events (LEs) of 139K users and 346K tracks collected from Twitter. The dataset comes with a rich set of item content features and user context features, and the timestamps of the LEs. Moreover, some of the user context features imply the cultural origin of the users, and some others&mdash;like hashtags&mdash;give clues to the emotional state of a user underlying an LE. In this paper, we provide some statistics to give insight into the dataset, and some directions in which the dataset can be used for making music recommendation. We also provide standardized training and test sets for experimentation, and some baseline results obtained by using factorization machines.</p> <p>The dataset contains three files:</p> <ul> <li>user_track_hashtag_timestamp.csv contains basic information about each listening event. For each listening event, we provide an id, the user_id, track_id, hashtag, created_at&nbsp;</li> <li>context_content_features.csv: contains all context and content features. For each listening event, we provide the id of the event, user_id, track_id, artist_id, content features regarding the track mentioned in the event (instrumentalness, liveness, speechiness, danceability, valence, loudness, tempo, acousticness, energy, mode, key) and context features regarding the listening event (coordinates (as geoJSON), place (as geoJSON), geo (as geoJSON), tweet_language, created_at, user_lang, time_zone, entities contained in the tweet).</li> <li>sentiment_values.csv contains sentiment information for hashtags. It contains the hashtag itself and the sentiment values gathered via four different sentiment dictionaries: AFINN, Opinion Lexicon, Sentistrength Lexicon and vader. For each of these dictionaries we list the minimum, maximum, sum and average of all&nbsp;sentiments of the tokens of the hashtag (if available, else we list empty values). However, as most hashtags only consist of a single token, these&nbsp;values are equal in most cases. Please note that the lexica are rather diverse and therefore, are able to resolve very different terms against a score. Hence,&nbsp;the resulting csv is rather sparse. The file contains the following comma-separated values: &lt;hashtag, vader_min, vader_max, vader_sum,vader_avg, &nbsp;afinn_min, afinn_max,&nbsp;afinn_sum, afinn_avg, ol_min, ol_max, ol_sum, ol_avg, ss_min, ss_max, ss_sum, ss_avg &gt;, where we abbreviate all scores gathered over the Opinion Lexicon with the&nbsp;prefix &#39;ol&#39;. Similarly, &#39;ss&#39; stands for SentiStrength.&nbsp;</li> </ul> <p>Please also find the training and test-splits for the dataset in this repo. Also, prototypical implementations of a context-aware recommender system based on the dataset can be found at&nbsp; <a href="https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM">https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM</a>.</p> <p>If you make use of this dataset, please cite the following paper where we describe and experiment with the dataset:</p> <p>@inproceedings{smc18,<br> title = {#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems},<br> author = {Asmita Poddar and Eva Zangerle and Yi-Hsuan Yang},<br> url = {http://mac.citi.sinica.edu.tw/~yang/pub/poddar18smc.pdf},<br> year = {2018},<br> date = {2018-07-04},<br> booktitle = {Proceedings of the 15th Sound &amp; Music Computing Conference},<br> address = {Limassol, Cyprus},<br> note = {code at https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM},<br> tppubtype = {inproceedings}<br> }</p>

opencc-by-4.0Jul 2018View details →
zenodo40/100

A small dataset for demonstrating the benchmarking of spot-detection/spot-counting workflows with BIAFLOWS

<p>The images were generated by&nbsp;<a href="http://www.cs.tut.fi/sgn/csb/simcep/tool.html">SIMCEP</a>, a widefield fluorescence microscopy biological images simulator.</p> <p>The dataset contains 5 input images and 5 ground-truth images with the suffix _lbl.</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

An in-situ daily dataset for benchmarking temporal variability of groundwater recharge

<p>A newly developed benchmark dataset of groundwater recharge per unit specific yield (RpSy, n meters) at daily temporal resolution is presented. The data has been obtained through the application of the Water table Fluctuation (WTF) method at groundwater wells within the continental US. To ensure high-fidelity estimates, only wells that meet a set of stringent criteria have been considered. The RpSy dataset may serve as a benchmark for validating the temporal consistency of recharge products and daily simulation results from land surface and integrated hydrologic models.</p> <p>The resulting product is a continuous daily RpSy (n meters) time series data for 485 groundwater wells. The data files are provided in the .csv format and consist of three columns for each observation well. The first column lists the local time, while the second and third columns provide the RpSy and RpSyu (considering a groundwater depth-dependent specific yield) time series in meters per day. Additionally, a file containing site information for all the selected wells is included. It contains four columns that detail the USGS ID of the groundwater well, its latitude (Lat), longitude (Long), and screen depth (depth, in meters). The data file can be accessed in most text editors and spreadsheets.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Dataset: Benchmark Problems for Simulating Hyperloop Aerodynamics

<p>Dataset for the results contained within 'Benchmarks problems for Simulating Hyperloop Aerodynamics', Lang et al., Phys. Fluids 36 (2024). doi.org/10.1063/5.0229914</p> <p>In this study, 3 benchmark problems for simulating the aerodynamics of a Hyperloop system are proposed.&nbsp;This dataset gives the raw data used to generate the figures and also coordinates of the geometries used in the simulations.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Dataset for Benchmarking the Sim-to-Real Gap in Cloth Manipulation

<p>This dataset&nbsp;is supplemental to the paper "Benchmarking the Sim-to-Real Gap in Cloth Manipulation".</p> <p>D. Blanco-Mulero, O. Barbany, G. Alcan, A. Colom&eacute;, C. Torras and V. Kyrki, "Benchmarking the Sim-to-Real Gap in Cloth Manipulation," in IEEE Robotics and Automation Letters, vol. 9, no. 3, pp. 2981-2988, March 2024, doi: 10.1109/LRA.2024.3360814</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

RSYD-BASIC results for AMR benchmarking dataset subset (MiSeq data)

<p><strong>Input data:</strong></p> <ul> <li>20240905_test_config.yaml: original config file used to run the pipeline</li> <li>20241004_rsyd_largeset_reads.zip: renamed Illumina MiSeq reads</li> <li>20241011-sample-overview.xlsx: overview of SRR accession numbers to internal sample numbers</li> <li>ILM_Run0001_Y20240904_kts_new.xlsx: runsheet&nbsp;</li> <li>input_en.yaml: column name configuration for the run</li> <li>lis_data.zip: LIS report and bacteria list used for LIS-specific results</li> </ul> <p><strong>Expected results:</strong></p> <ul> <li>20240910_test_illumina_largeset.zip: Results of the RSYD-BASIC pipeline, version 1.15.1, with the reads used</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Birdsong NOIZEUS: Bioacoustics noise reduction benchmark dataset

<p>------------------------------------------------------------------------<br>Birdsong noizeus dataset<br>------------------------------------------------------------------------</p> <p>Authors: Tim Sainburg &amp; Asaf Zorea<br>Year: 2024</p> <p>------------------------------------------------------------------------<br>General information<br>------------------------------------------------------------------------<br>- There are 5 recordings from each of 14 individuals (European starlings) recorded in an acoustically isolated chamber.&nbsp;<br>- For each song, we apply noise at 5 SNR levels (0dB, 5dB, 10dB, 15dB).&nbsp;<br>- There are 8 noise types, each taken from a single noise clip from the "Soundscapes from around the world" dataset.<br>- They are "rain", "town", "wind", "waterfall", "insect", "swamp" "frogscape", "forest"<br>- Each soundscape contains multiple noise sources.&nbsp;<br>- Noise levels were estimated using the pyloudnorm software (Steinmetz et al., 2021)<br>- Audio is provided as waveforms at 44100 samplerate</p> <p>------------------------------------------------------------------------<br>Data format<br>------------------------------------------------------------------------</p> <p>- clean<br>&nbsp; &nbsp; - {bird_name}_{timestamp}.wav<br>- noisy<br>&nbsp; &nbsp; - {snr}dB<br>&nbsp; &nbsp; &nbsp; &nbsp; - {bird_name}_{timestamp}_{noise_category}_{snr}.wav<br>- noise_sample<br>&nbsp; &nbsp; - {snr}dB<br>&nbsp; &nbsp; &nbsp; &nbsp; - {bird_name}_{timestamp}_{noise_category}_{snr}.wav</p> <p>`clean` contains the original clean audio.<br>`noisy` contains the song+noise<br>`noise_sample` contains a 1-second sample of noise only.&nbsp;</p> <p>Timestamp is in the format YYYY-MM-DD_HH-MM-SS-MILLISECONDS and refers to the time that the song was recorded.&nbsp;</p> <p>------------------------------------------------------------------------<br>Data sources<br>------------------------------------------------------------------------</p> <p>Birdsong<br>---------<br>Birdsong are acoustically isolated songs from 14 European Starlings<br>https://zenodo.org/records/3237218</p> <p>Citation:&nbsp;<br>Arneodo, Z., Sainburg, T., Jeanne, J., &amp; Gentner, T. (2019). An acoustically isolated European starling song library (Version v1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.3237218</p> <p>This dataset is available under the following license:<br>- &nbsp;Creative Commons Attribution 4.0 International (https://creativecommons.org/licenses/by/4.0/legalcode)</p> <p><br>Noise<br>-----<br>Noise are taken from the Xeno-canto - "Soundscapes from around the world" dataset<br>https://www.gbif.org/dataset/ff571aeb-46bf-45c4-ad2c-af4d68315765</p> <p>Citation:<br>Vellinga W (2024). Xeno-canto - Soundscapes from around the world. Xeno-canto Foundation for Nature Sounds. Occurrence dataset https://doi.org/10.15468/9u3zaq accessed via GBIF.org on 2024-10-17.</p> <p>We sampled 8 soundscapes from this dataset:<br>&nbsp; &nbsp; 1. Rain https://www.gbif.org/occurrence/4523646364<br>&nbsp; &nbsp; &nbsp; &nbsp; - Virginia<br>&nbsp; &nbsp; &nbsp; &nbsp; - 457s<br>&nbsp; &nbsp; &nbsp; &nbsp; - rain, Recording of the feeders in the backyard with a light rain falling on the leaf litter, while a freight train passes by ~1/2 mile away.<br>&nbsp; &nbsp; 2. Town https://xeno-canto.org/696263<br>&nbsp; &nbsp; &nbsp; &nbsp; - 2:22<br>&nbsp; &nbsp; &nbsp; &nbsp; - &nbsp;the closer habitat are greater trees and coniferes ..smaller bushes ...some other Krautg&auml;rten ... traffic-noise is to hear..as airplanes,too.<br>&nbsp; &nbsp; &nbsp; &nbsp; - insects, birds, bells, traffic?, airplane<br>&nbsp; &nbsp; &nbsp; &nbsp; - Germany<br>&nbsp; &nbsp; 3. Wind https://xeno-canto.org/911773<br>&nbsp; &nbsp; &nbsp; &nbsp; - 5:29<br>&nbsp; &nbsp; &nbsp; &nbsp; - Very windy day, grassland to shrubland habitat<br>&nbsp; &nbsp; &nbsp; &nbsp; - Wisconsin<br>&nbsp; &nbsp; 4. Waterfall https://xeno-canto.org/406993<br>&nbsp; &nbsp; &nbsp; &nbsp; - 24:02<br>&nbsp; &nbsp; &nbsp; &nbsp; - sound scenes captured in a river forest with zarzas, hiedras and other matorrales at the edge of a small waterfall.<br>&nbsp; &nbsp; &nbsp; &nbsp; - Spain<br>&nbsp; &nbsp; 5. Insect https://xeno-canto.org/454914<br>&nbsp; &nbsp; &nbsp; &nbsp; - 5:45<br>&nbsp; &nbsp; &nbsp; &nbsp; - Australia<br>&nbsp; &nbsp; &nbsp; &nbsp; - cicadias, birds,<br>&nbsp; &nbsp; 6. Swanp https://xeno-canto.org/909875<br>&nbsp; &nbsp; &nbsp; &nbsp; - 2:24<br>&nbsp; &nbsp; &nbsp; &nbsp; - Swampy area along dirt road/trail. Species Include: American Bullfrog, Cricket Frogs, American Crow, Prothonotary Warbler, Red-winged Blackbird, Indigo Bunting, Northern Cardinal<br>&nbsp; &nbsp; 7. Frogscape https://xeno-canto.org/718213<br>&nbsp; &nbsp; &nbsp; &nbsp; - 4:42&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; - European Tree Frog Hyla arborea Green Frog Pelophlyax sp.<br>&nbsp; &nbsp; &nbsp; &nbsp; - Austrial&nbsp;<br>&nbsp; &nbsp; 8. Forest https://xeno-canto.org/900744&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; - 3:04<br>&nbsp; &nbsp; &nbsp; &nbsp; - Sweden<br>&nbsp; &nbsp; &nbsp; &nbsp; - A beautiful choir with frogs, toads, geese and ducks to enjoy in the darkness a foggy night.<br>&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;<br>All noise recordings have one of the following licensesL<br>- Creative Commons Attribution-NonCommercial-ShareAlike 4.0 (https://creativecommons.org/licenses/by-nc-sa/4.0/)<br>- Creative Commons Attribution-ShareAlike 4.0 (https://creativecommons.org/licenses/by-sa/4.0/)</p> <p><br>Additional information about the noise dataset</p> <p><br>------------------------------------------------------------------------<br>citations<br>------------------------------------------------------------------------</p> <p>Steinmetz, C. J., &amp; Reiss, J. (2021, May). pyloudnorm: A simple yet flexible loudness meter in python. In Audio Engineering Society Convention 150. Audio Engineering Society.</p> <p>Arneodo, Z., Sainburg, T., Jeanne, J., &amp; Gentner, T. (2019). An acoustically isolated European starling song library (Version v1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.3237218</p> <p>Vellinga W (2024). Xeno-canto - Soundscapes from around the world. Xeno-canto Foundation for Nature Sounds. Occurrence dataset https://doi.org/10.15468/9u3zaq accessed via GBIF.org on 2024-10-17.</p>

opencc-by-nc-4.0Oct 2024View details →
zenodo40/100

Ice Anatomy: A Benchmark Dataset and Methodology for Automatic Ice Boundary Extraction from Radio-Echo Sounding Data

<p>The measurement of ice thickness is of great importance for the accurate estimation of glacier volume and the delineation of their bedrock topography. In particular, this is a crucial factor in forecasting the future evolution of glaciers in the context of a changing climate. In order to derive the ice thickness, the travel time of electromagnetic waves in radargrams acquired by radio-echo sounding (RES) systems is analyzed. This can only be achieved by identifying the ice surface and underlying ice bottom in corresponding radargrams. Manually identifying these two reflection horizons in RES data is a laborious and time-consuming process. Consequently, scientists are attempting to automate this task through the use of techniques such as deep learning. Such automation can significantly reduce the time between a field campaign and the calculation of the glacier's ice thickness distribution. In this paper, we present the first benchmark dataset for delineating the ice surface and bottom boundaries in RES data, to facilitate straightforward comparisons of deep learning models in the future. The ``IceAnatomy'' dataset comprises radargrams and the corresponding manual picks, amounting to a total of over 45,000km of observations. The RES data originates from three sources: FAU, CReSIS, and AWI. The dataset comprises different RES systems as well as different pre-processing methods. In addition, the data was acquired over a large range of geographical and glaciological settings, featuring different thermal regimes present in Antarctica and the Southern Patagonian Icefield. This diversity ensures that the models' behaviors can be analyzed in different scenarios. We define a standardized train-test split for each source in the dataset. This allows us to introduce not only a baseline model trained on the entire training set (the ``omni'' model), but also three source-specific baseline models. The source-specific models are trained exclusively on the subset of the training data acquired by the specified source. The baseline models provide an initial benchmark against which subsequent models can be compared. The source-specific models demonstrate more accurate results than the omni model. For the FAU, CReSIS, and AWI test sets, the source-specific models achieve &nbsp;Mean Meter Errors of 2.1m, 23.1m, and 4.9m for the ice surface and 9.1m, 78.2m, and 29.3m for the ice bottom. In relation to the mean measured ice thickness, these errors equate to 1.2%, 3.1%, and 0.3% for the ice surface and&nbsp; 4.9%, 10.4%, and 1.5% for the ice bottom.</p> <p>&nbsp;</p> <p>&nbsp;For more information, please read the following paper:</p> <p>[Coming soon. Currently under review.]</p> <p>Please also cite this paper if you plan on using the dataset.</p> <p>&nbsp;</p> <p>For the implementation of a baseline model please visit:</p> <p>[Coming soon]</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

LiDAR Segmentation Benchmark Dataset: La Palma

<p>The purpose of this benchmark dataset is primarily to allow the assessment of individual tree segmentation algorithms on LiDAR data. The high density ULS point cloud data were captured by the Zenmuse L1 sensor on a DJI Matrice 300RTK over a coniferous plot on La Palma (Canary Islands, Spain). The dataset comprises the CHM and the normalized 3D cloud, as well as the fully annotated data to serve as a benchmark for segmentation.</p> <p>The publication associated with this database is:</p> <p>Marcello, J.; Sp&iacute;nola, M.; Albors, L.; Marqu&eacute;s, F.; Rodr&iacute;guez-Esparrag&oacute;n, D.; Eugenio, F. Performance of Individual Tree Segmentation Algorithms in Forest Ecosystems Using UAV LiDAR Data. <em>Drones</em>&nbsp;<strong>2024</strong>,&nbsp;<em>8</em>, 772. https://doi.org/10.3390/drones8120772</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

WONDERBREAD: A Benchmark + Dataset for Business Process Management (BPM) Tasks

<p><strong>Paper:</strong> <a href="https://arxiv.org/abs/2406.13264">WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks</a></p> <h2><strong>Background</strong></h2> <p>The&nbsp;<em>WONDERBREAD</em> dataset contains <strong>2,928 human demonstrations</strong> of <strong>598 web navigation workflows</strong> across <strong>6 types of BPM tasks</strong>. These tasks measure the ability of a model to generate accurate documentation, assist in knowledge transfer, and improve the efficiency of workflows.</p> <p>Please see our website for more details:&nbsp;<a href="https://wonderbread.stanford.edu/">https://wonderbread.stanford.edu/</a></p> <h2><strong>Quick Start</strong></h2> <p>To start, download <strong>debug_demos.zip</strong> (1 GB). It contains a subset of <strong>24 demonstrations</strong> which can give you a sense of how the dataset is structured.</p> <p>To reproduce the paper, download <strong>gold_demos.zip</strong> (33 GB). It contains <strong>724 demonstrations</strong> corresponding to the 162 "Gold" tasks which were used for all the evaluations in the original paper.</p> <p>To obtain the full dataset, download&nbsp;<strong>demos.zip</strong> (133 GB). This contains all <strong>2,928 demonstrations</strong> and can be used for training, fine-tuning, and evaluating models.</p> <h2><strong>Dataset Structure</strong></h2> <p>The dataset contains several files, defined below.</p> <ol> <li><strong>Raw Data</strong><em> (useful for training/fine-tuning/evaluation)</em> <ol> <li><strong>debug_demos.zip </strong>(1 GB)<strong> -- </strong>a subset of only 24 demonstrations taken from the full dataset. Useful to get a sense of the dataset and for debugging.</li> <li><strong>gold_demos.zip</strong> (22 GB) -- a subset of only 724 demonstrations corresopnding to the 162 "Gold" tasks. This is the dataset that was used for all evaluations in the original <em>WONDERBREAD</em> paper.</li> <li><strong>demos.zip</strong> (133 GB) -- all 2,928 demonstrations across 598 tasks. Useful for training your own models.</li> </ol> </li> <li><strong>Modality-Specific Subsets of Raw Data </strong><em>(useful for specific types of training/fine-tuning/evaluation)</em><br> <ol> <li><strong>All Demos</strong> <ol> <li><strong>demos_sop_only.zip</strong> (4 MB)-- only the SOP <code>.txt</code> files for all 2,928 demonstrations</li> <li><strong>demos_sop_and_trace_only.zip</strong> (770 MB)-- only the SOP <code>.txt</code> files and action trace <code>.json</code> files for all 2,928 demonstrations</li> <li><strong>demos_sop_and_trace_and_screenshots_only.zip</strong> (22 GB)-- only the SOP <code>.txt</code> files and action trace <code>.json</code> files and screenshot images for all 2,928 demonstrations</li> </ol> </li> <li><strong>"Gold" Demos</strong> <ol> <li><strong>gold_demos_sop_only.zip</strong> (1 MB)-- only the SOP <code>.txt</code> files for the 724 demonstrations in the "Gold" tasks.</li> <li><strong>gold_demos_sop_and_trace_only.zip</strong> (190 MB) -- only the SOP <code>.txt</code> files and action trace&nbsp;<code>.json</code>&nbsp;files for the 724 demonstrations in the "Gold" tasks.</li> <li><strong>gold_demos_sop_and_trace_and_screenshots_only.zip </strong>(6 GB) -- only the SOP <code>.txt</code> files and action trace <code>.json</code> files and screenshot images for the 724 demonstrations in the "Gold" tasks</li> </ol> </li> </ol> </li> <li><strong>Evaluation</strong><em> (useful for evaluation)</em> <ol> <li><strong>qa_dataset.csv -- </strong>contains all 120 questions and ground truth answers used in the "Knowlege Transfer" evaluation.<strong><br></strong></li> <li><strong>df_rankings.csv -- </strong>contains the rankings of all "Gold" tasks used in the "SOP Ranking" evaluation.<strong><br></strong></li> </ol> </li> <li><strong>Metadata</strong><em> (can be safely ignored)</em> <ol> <li><strong>Process Mining Task Demonstrations.xlsx --</strong> maps human annotators to specific demonstrations; also contains "Gold" task rankings used in the "SOP Ranking" evaluation.</li> <li><strong>metadata.json -- </strong>maps Google Drive URLs to Google Drive Folder IDs to demonstration names</li> <li><strong>df_valid.csv -- </strong>tracks assets associated with each demonstration</li> </ol> </li> </ol>

opencc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record