Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,549
datasets available to search
ShareScore release 0.9.0
Dataset results
1,549 results for “benchmarks”
Figure 8 in Unraveling distributional patterns and life-history traits of a deep-water shrimp Plesionika edwardsii (Decapoda, Pandalidae) under unexploited virgin conditions: a benchmark for fisheries management
Figure 8. Hypothesized life cycle of Plesionika edwardsii in the Azorean region. After the incubation period of shrimp eggs, (1) larvae are released into the water column and (2) juveniles develop in shallow waters. Mature females and males are distributed up to 600 m with a sexual segregation by depth: (3) non-ovigerous females are mainly found up to 200 m, (4) ovigerous females between 200 and 300 m, and (5) males from 400 to 500 m deep. Females are bigger than males, and ovigerous females are bigger than nonovigerous females. A bigger-deeper trend is observed up to 400 m. (6) Long larval stages of P. edwardsii increases its potential for dispersal (Landeira et al., 2009), favoring connectivity and stock homogeneity between adjacent areas.
Figure 5 in Unraveling distributional patterns and life-history traits of a deep-water shrimp Plesionika edwardsii (Decapoda, Pandalidae) under unexploited virgin conditions: a benchmark for fisheries management
Figure 5. Sex ratio of Plesionika edwardsii by depth stratum in the Azorean region during the period 1999–2000.
Figure 2 in Unraveling distributional patterns and life-history traits of a deep-water shrimp Plesionika edwardsii (Decapoda, Pandalidae) under unexploited virgin conditions: a benchmark for fisheries management
Figure 2. Seasonal predicted mean catch per unit effort (CPUE, g trap-1) by depth stratum for males, non-ovigerous and ovigerous females of Plesionika edwardsii in the Azorean region for the period 1999–2000. Light-colored symbols represent raw data. Detailed parameter estimates are in Tab. S4.
Figure 7 in Unraveling distributional patterns and life-history traits of a deep-water shrimp Plesionika edwardsii (Decapoda, Pandalidae) under unexploited virgin conditions: a benchmark for fisheries management
Figure 7. Size at which 50 % of the shrimps are mature (L 50) estimated for Plesionika edwardsii in the Azorean region fitting a logistic curve to the proportion of ovigerous females. Logistic curve was estimated combining all data obtained during the period 1999–2000.
Figure 4 in Unraveling distributional patterns and life-history traits of a deep-water shrimp Plesionika edwardsii (Decapoda, Pandalidae) under unexploited virgin conditions: a benchmark for fisheries management
Figure 4. Seasonal predicted mean cephalothorax length (CL) by depth stratum for males, non-ovigerous and ovigerous females of Plesionika edwardsii in the Azorean region for the period 1999–2000. Light-colored symbols represent raw data. Detailed parameter estimates are in Tab. S4.
Figure 1 in Unraveling distributional patterns and life-history traits of a deep-water shrimp Plesionika edwardsii (Decapoda, Pandalidae) under unexploited virgin conditions: a benchmark for fisheries management
Figure 1. Sampling areas of Plesionika edwardsii in the mid-North Atlantic Ocean, Azorean region (ICES Subdivision 10a2) between 1999 and 2000. Orange dots represent each site sampled by a trap.
Figure 6 in Unraveling distributional patterns and life-history traits of a deep-water shrimp Plesionika edwardsii (Decapoda, Pandalidae) under unexploited virgin conditions: a benchmark for fisheries management
Figure 6. Sex ratio of Plesionika edwardsii by size class in the Azorean region during the period 1999–2000.
Figure 3 in Unraveling distributional patterns and life-history traits of a deep-water shrimp Plesionika edwardsii (Decapoda, Pandalidae) under unexploited virgin conditions: a benchmark for fisheries management
Figure 3. Size frequency distribution of males, non-ovigerous and ovigerous females Plesionika edwardsii in the Azorean region during the period 1999-2000.
Processed Datasets - Imputation in Well Log Data: A Benchmark
<p>Imputation of well log data is a common task in the field. However a quick review of the literature reveals a lack of padronization when evaluating methods for the problem. The goal of the benchmark is to introduce a standard evaluation protocol to any imputation method for well log data. </p> <p>In the proposed benchmark, three public datasets are used:</p> <ul> <li><strong>Geolink:</strong> The Geolink Dataset is another public dataset of wells in the Norwegian offshore. The data is provided by the company of the same name, <a href="https://www.geolink-s2.com/" target="_blank" rel="noopener">GEOLINK</a> and follows the NOLD 2.0 license. <br>This dataset contains a total of 223 wells. It also has lithology labels for the wells with a total of 36 lithology classes. [<a href="https://drive.google.com/drive/folders/1EgDN57LDuvlZAwr5-eHWB5CTJ7K9HpDP" target="_blank" rel="noopener">download original</a>]</li> <li><strong>Taranaki Basin:</strong> The Taranaki Basin Dataset is a curated set of wells and a convenient option for experimentation especially due to it is ease of accessibility and use.<br>This collection, under the CDLA-Sharing-1.0 license, contains well logs extracted from the <a href="https://geodata.nzpam.govt.nz/" target="_blank" rel="noopener">New Zealand Petroleum & Minerals Online Exploration Database</a> and <a href="http://pet.gns.cri.nz/" target="_blank" rel="noopener">Petlab</a>.<br>There are a total of 407 wells, of which 289 are onshore and 118 are offshore exploration and production wells. [<a href="https://developer.ibm.com/exchanges/data/all/taranaki-basin-curated-well-logs/" target="_blank" rel="noopener">download original</a>]</li> <li><strong>Teapot Dome:</strong> The Teapot Dome dataset is provided by the Rocky Mountain Oilfield Testing Center (RMOTC) and the US Department of Energy.<br>It contains different types of data related to the Teapot Dome oil field, such as 2D and 3D seismic data, well logs, and GIS data. The data is licensed under the Creative Commons 4.0 license. <br>In total, the dataset has 1,179 wells with available logs. The number of available logs varies across wells. There are only 91 wells with the gamma ray, bulk density, and neutron porosity logs, while only three wells have the complete basic suite. [<a href="http://s3.amazonaws.com/open.source.geoscience/open_data/teapot/rmotc.tar" target="_blank" rel="noopener">direct download</a>]</li> </ul> <p>Here you can download all three datasets already preprocessed to be used with our implementation, found <a href="https://github.com/uai-ufmg/well-log-imputation" target="_blank" rel="noopener">here</a>.</p> <p> </p> <h3>File Description:</h3> <p>There are six files for each fold partition for each dataset.</p> <ul> <li><code><em>datasetname_fold_k_well_log_metadata_train.json </em></code>: JSON file with general information of the slices of <strong>training </strong>partition of the fold <strong>k</strong>. Contains total number of slices and the number of slices per well.<em> </em></li> <li><em><code>datasetname_fold_k_well_log_metadata_val.json</code> </em>: JSON file with general information of the slices of <strong>validation </strong>partition of the fold <strong>k</strong>. Contains total number of slices and the number of slices per well. </li> <li><em><code>datasetname_fold_k_well_log_slices_train.npy</code>: </em>.npy (numpy) file ready to be loaded with the slices for <strong>training </strong>of the fold <strong>k </strong>already processed. When loaded<em> </em>should have shape of<em> (total_slices, 256, number_of_logs)</em></li> <li><em><code>datasetname_fold_k_well_log_slices_val.npy</code> </em>: .npy (numpy) file ready to be loaded with the slices for <strong>validation </strong>of the fold <strong>k </strong>already processed.</li> <li><em><code>datasetname_fold_k_well_log_slices_meta_train.json</code> : </em>JSON file with the slices info for all slices in the <strong>training </strong>partition of the fold <strong>k</strong>. For each slice, 7 data points are provided, the last four are discarded (it would contain other information that was not used). The first three are in order the: origin well name, the starting position in that well, and the end position of the slice in that well.</li> <li><em><code>datasetname_fold_k_well_log_slices_meta_val.json</code> </em>: JSON file with the slices info for all slices in the <strong>validation </strong>partition of the fold <strong>k</strong>.</li> </ul>
tBiomed: Semantic Table Annotations Benchmark for Biomedical Domain
<p><strong>tBiomed </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p><strong>tBiomed </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using two levels of a recursive hierarchy of related concepts in Wikidata.</p> <p><strong>tBiomed </strong>contains <strong>26,778</strong> entity and horizontal tables, while this repository contains only a <strong>validation fold</strong> of the original data representing <strong>20%</strong> of the total of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>1</strong> <strong>GB</strong>.</p> <p>We included the full version of the dataset. We will update this repository ground truth data of the test set in the Future.</p> <p>The supported tasks for semantic table annotations are: </p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> <li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li> </ol>
Context Map for eshopContainers Benchmark
<p>This is part of the outcomes from the paper titled "Deriving Microservice Architectural Perspectives Using Static Code Analysis For C# Platform"</p>
LC-MS/MS, PSM and BLG properties data from "Benchmarking the identification of a single degraded protein to explore optimal search strategies for ancient proteins"
<p>This dataset contains the data analyzed in:</p> <p>Rodriguez Palomo I, Nair B, Chang Y, Dartigues B, Dekker K, Mackie M, Evans M, Macleod R, Olsen JV, Collins MJ. (2023) "<em>Benchmarking the identification of a single degraded protein to explore optimal search strategies for ancient proteins"</em></p> <p>It contains the following data:</p> <ul> <li>raw_files.zip Thermo RAW files for the 0, 4 and 128 days samples</li> <li>benchmark_results.zip PSMs data from the analysis of the RAW files <ul> <li>Data from runs in Mascot, Fragpipe, pFind, Metamorpheus, MaxQuant and DeNovoGUI</li> <li>Parameters and workflow files for MaxQuant and Fragpipe</li> </ul> </li> <li>bovin_blg_prop.zip BLG properties files: amyloid formation, 3D structure and solvent accessibility</li> <li>benchmark_table.csv Table with runs settings for benchmarking</li> <li>parameters_table.xlsx Spreadsheet with software parameters, derived from files used to run each software</li> </ul> <p> </p>
Supporting data for "Benchmarking the integration of hexagonal boron nitride crystals and thin films into graphene-based van der Waals heterostructures"
<p>Dataset for the publication "Benchmarking the integration of hexagonal boron nitride crystals and thin films into graphene-based van der Waals heterostructures"</p>
A software benchmark for cardiac elastodynamics
<p>Data used in the article:<br>"A software benchmark for cardiac elastodynamics" by Arostica et al, Computer Methods in Applied Mechanics and Engineering. The DOI of the article was not available at the moment of publishing this data set.</p>
Silicodata: An Annotated Benchmark CXR Dataset for Silicosis Detection
Open the record for dataset details and reuse information.
Simulation data for benchmarking de novo long read transcriptome assembly software
<p>Method of simulation of differentially expressed biological replicates</p> <p>We first obtained a subset of transcripts that are widely expressed in the GTEx v9 dataset (92 samples) using Gencode comprehensive annotation (v44). We kept transcripts with more than 5 reads in at least 15 samples after Salmon quantification (18145 genes, 40509 transcripts), and stored their mean count per million (CPM) values as the control group’s baseline expression. We then generated a perturbed set of CPM values where transcript expression was changed by: (1) randomly selecting 1000 genes and changing all transcripts belonging to that gene concordantly (500 genes 2 fold up and 500 genes 2 fold down), (2) selected another 1000 genes randomly, and then select 2 random transcripts from the gene and swap their expression, (3) selected another 1000 genes randomly, and then select 1 random transcript to change its expression (500 transcripts 2 fold up and 500 transcripts 2 fold down). The updated CPM were stored as the perturbed group baseline expression. We then generated a count matrix and CPM matrix for 3 control replicates and 3 perturbed replicates with gamma distribution, followed by a Poisson distribution <a href="https://www.zotero.org/google-docs/?cUP4ui">(Baldoni et al., 2024)</a>. Both long-read and short-read FASTQ files were simulated using SQANTI-SIM with default settings and ONT R9.4 cDNA error profile (v 0.2.1) <a href="https://www.zotero.org/google-docs/?Qyopst">(Mestre-Tomás et al., 2023)</a>. The long read data contained 6 million reads in total, and an average read length of 1085 bp, and short read data was 100 bp paired-end. We then subsampled the short-read data to match the total number of base pairs in the long read data (6.5 billion bases). The simulated data was non-stranded, and contains 2000 DE genes, 2000 genes with DTU, 5927 transcripts with DTU and 6933 DE transcripts.</p> <p> </p>
Evaluation of Spatiotemporal Fusion Methods Using Sentinel-2 And Sentinel-3: A New Benchmark Dataset And Comparison
<p>In Earth observation, data fusion is important to generate high temporal and spatial resolution images. Nevertheless, existing research on data fusion primarily concentrates on merging two sources of data (mostly MODIS and Landsat). Therefore, we offer the community a new benchmark dataset for evaluating data fusion using new European sensors (Sentinel-2 and Sentinel-3).</p> <p>The dataset is composed of three different sites located in different parts of the world to ensure the diversity of the ecosystem. The two components of the dataset are collected from operating missions ( Sentinel-2 and Sentinel-3). We also provide 10 bands for Sentinel-2 ranging from blue to SWIR, 4 bands at 10m resolution and 6 at 20m resolution. For Sentinel-3 16 bands are provided with a spatial resolution of 300m. The multiple bands allow for different applications for this dataset such as testing data fusion methods, etc.</p>
A benchmark dataset for the grazing flow over porous materials
<p>Wind-tunnel data of a grazing flow over porous wall-inserts to be used as a benchmark dataset for the development and validation of numerical modeling approaches of flows over and through porous media. </p> <p>The dataset contains several profiles along the streamwise extent of the wall-insert of the mean velocity magnitude, the turbulent intensity, and the turbulent length scale. Also included are some boundary-layer parameters of these profiles. These have been derived from single-component constant temperature hot-wire measurements.</p> <p>Additionally, the spectra of the unsteady wall-pressure fluctuations at several locations on the upper and lower surfaces of the porous wall-inserts are provided. These unsteady pressure measurements have been acquired using semi-infinite waveguide-type remote-microphone probes.</p> <p>Tested are two porous media with the same <em>diamond-lattice</em> pattern structure but different permeabilities and a reference solid-walled case. All three cases are tested at three inflow velocities: 15 m/s, 20 m/s, and 25 m/s.</p> <p> </p> <p>Modification in v3: Correction of permeability values in Table 1 on page 3 of <em>AIAA_Manuscript_GrazingFlowPorousMaterials_v3.pdf</em>.</p>
GloMPO (Globally Managed Parallel Optimization) Benchmark Test Data
<p>Dataset associated with:<br> M. Freitas Gustavo and T. Verstraelen (2021), "GloMPO (Globally Managed Parallel Optimization) - a tool for expensive, black-box optimizations: application to ReaxFF reparameterizations".</p> <p>Contains optimization trajectories using the CMA-ES optimizer applied to various benchmark functions and ReaxFF reparameterizations. Compares optimization results using the GloMPO framework (github.com/mfgustavo/glompo) to unmanaged optimization results.</p>
Alignment files for coverage benchmarks: Illumina and Nanopore sequencing datasets
<ul> <li><strong>cpara-illumina-noseq.bam</strong> and <strong>cpara-ont-noseq.bam</strong>: BAM files produced aligning the raw reads produced respectively by Illumina NextSeq and ONT Nanopore sequencing of an isolate of <em>C. parapsilosis</em> to evaluate the coverage calculations using real datasets.*</li> <li><strong>HG00258.bam</strong>: Exome sequencing from the 1000 Genomes Project (Clarke et al 2016 <a href="https://doi.org/10.1093/nar/gkw829">https://doi.org/10.1093/nar/gkw829</a>).</li> <li><strong>panel_01.bam</strong>: targeted sequencing of a Human gene panel of 16 genes.*</li> </ul> <p>* Sequences and qualities have been removed</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.