Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,549

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,549 results for β€œbenchmarks”

Learn how ShareScore rates datasets β†—
zenodo44/100

PheKnowLator Human Disease KG Benchmarks: Class-Standard Relations-OWLNETS (v2.0.0 - January 2021)

<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds&nbsp;(v2.0.0)</strong></p><p><strong>Build Type:&nbsp;</strong><i>Class-Standard Relations-OWLNETS</i></p><p><strong>Build Date: </strong>January 25, 2021</p><p>&nbsp;</p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p>&nbsp;</p><p>🚨&nbsp;<strong>AVAILABLE FILES&nbsp;</strong>🚨&nbsp;</p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page&nbsp;πŸ‘‰&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/January-25%2C-2021">here</a>.</li></ul>

opencc-by-4.0Jan 2021View details β†’
zenodo44/100

PheKnowLator Human Disease KG Benchmarks: Class-Inverse Relations-OWLNETS (v2.0.0 - January 2021)

<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds&nbsp;(v2.0.0)</strong></p><p><strong>Build Type:&nbsp;</strong><i>Class-Inverse Relations-OWLNETS</i></p><p><strong>Build Date: </strong>January 25, 2021</p><p>&nbsp;</p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p>&nbsp;</p><p>🚨&nbsp;<strong>AVAILABLE FILES&nbsp;</strong>🚨&nbsp;</p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page&nbsp;πŸ‘‰&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/January-25%2C-2021">here</a>.</li></ul>

opencc-by-4.0Jan 2021View details β†’
zenodo44/100

PheKnowLator Human Disease KG Benchmarks: Instance-Inverse Relations-OWL (v2.0.0 - January 2021)

<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds&nbsp;(v2.0.0)</strong></p><p><strong>Build Type:&nbsp;</strong><i>Instance-Inverse&nbsp;Relations-OWL</i></p><p><strong>Build Date: </strong>January 25, 2021</p><p>&nbsp;</p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p>&nbsp;</p><p>🚨&nbsp;<strong>AVAILABLE FILES&nbsp;</strong>🚨&nbsp;</p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page&nbsp;πŸ‘‰&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/January-25%2C-2021">here</a>.</li></ul>

opencc-by-4.0Jan 2012View details β†’
zenodo44/100

PheKnowLator Human Disease KG Benchmarks: Class-Inverse Relations-OWL (v2.0.0 - May 2020)

<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds&nbsp;(v2.0.0)</strong></p><p><strong>Build Type:&nbsp;</strong><i>Class-InverseRelations-OWL</i></p><p><strong>Build Date:&nbsp;</strong>May 10, 2020</p><p>&nbsp;</p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p>&nbsp;</p><p>🚨&nbsp;<strong>AVAILABLE FILES&nbsp;</strong>🚨&nbsp;</p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page&nbsp;πŸ‘‰&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/May-10%2C-2020">here</a>.</li></ul>

opencc-by-4.0Apr 2020View details β†’
zenodo44/100

PheKnowLator Human Disease KG Benchmarks: Class-Standard Relations-OWLNETS (v2.0.0 - May 2020)

<p><strong>PKT Human Disease Knowledge Graph Benchmark Builds&nbsp;(v2.0.0)</strong></p><p><strong>Build Type:&nbsp;</strong><i>Class-Standard Relations-OWLNETS</i></p><p><strong>Build Date:&nbsp;</strong>May 10, 2020</p><p>&nbsp;</p><h3><strong>Important Build Information</strong></h3><p>The benchmarks were originally built and stored using Google Cloud Platform (GCP) resources. For details and a complete description of this process, can be found on GitHub (<a href="https://github.com/callahantiff/PheKnowLator/tree/master/builds#readme">here</a>). Note that we have developed an archive for the builds on Zenodo. While the original GCP resources contained all associated files, due to the file size upload limits associated with each archive, we have limited the uploaded files to the KGs, associated metadata, and log files. The list of resources, including their URLs, and date of download, can all be found in the associated logs.</p><p>Details on each of the files generated by the build process can be found in the file associated with this directory (<a href="https://zenodo.org/records/10065431/files/PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx?download=1">PheKnowLator_HumanDiseaseKG_Output_FileInformation.xlsx</a>).</p><p>&nbsp;</p><p>🚨&nbsp;<strong>AVAILABLE FILES&nbsp;</strong>🚨&nbsp;</p><ul><li>Available KG benchmark files are zipped and listed below.</li><li>For additional details on what each file contains, please see the associated Wiki page&nbsp;πŸ‘‰&nbsp;<a href="https://github.com/callahantiff/PheKnowLator/wiki/May-10%2C-2020">here</a>.</li></ul>

opencc-by-4.0Apr 2020View details β†’
zenodo44/100

BUMP: A Benchmark of Reproducible Breaking Dependency Updates

<p>Bump is a benchmark of breaking dependency updates. A breaking update is defined as a pair of commits for a Java project, which we designate as the pre-commit and the breaking-commit. When we build the project with the pre-commit, compilation and test execution are successful, while the build of the breaking-commit fails. Each breaking-commit is a one-line change in the Maven pom file.</p>

openmit-licenseOct 2023View details β†’
zenodo44/100

Identifying patterns and recommendations of and for sustainable open data initiatives: a benchmarking-driven analysis of open government data initiatives among European countries

<p>This dataset contains data collected during a study <a href="https://www.sciencedirect.com/science/article/pii/S0740624X23000989"><em><strong>"Identifying patterns and recommendations of and for sustainable open data initiatives: a benchmarking-driven analysis of open government data initiatives among European countries"</strong></em></a> conducted by <em>Martin Lnenicka (University of Pardubice, Pardubice, Czech Republic), Anastasija Nikiforova (University of Tartu, Tartu, Estonia), Mariusz Luterek (University of Warsaw, Warsaw, Poland), Petar Milic (University of Pristina - Kosovska Mitrovica, Kosovska Mitrovica, Serbia), Daniel Rudmark (University of Gothenburg and RISE Research Institutes of Sweden, Gothenburg, Sweden), Sebastian Neumaier (St. P&ouml;lten University of Applied Sciences, Austria), Caterina Santoro (KU Leuven, Leuven, Belgium), Cesar Casiano Flores (University of Twente, Twente, the Netherlands), Marijn Janssen (Delft University of Technology, Delft, the Netherlands), Manuel Pedro Rodr&iacute;guez Bol&iacute;var (University of Granada, Granada, Spain).</em></p> <p>It is being made public both to act as supplementary data for "<em>Identifying patterns and recommendations of and for sustainable open data initiatives: a benchmarking-driven analysis of open government data initiatives among European countries</em>", Government Information Quarterly*, and in order for other researchers to use these data in their own work.&nbsp;</p> <p>***Methodology***</p> <p>The paper focuses on benchmarking of open data initiatives over the years and attempts to identify patterns observed among European countries that could lead to disparities in the development, growth, and sustainability of open data ecosystems.&nbsp;</p> <p>This study examines existing benchmarks, indices, and rankings of open (government) data initiatives to find the contexts by which these initiatives are shaped, both of which then outline a protocol to determine the patterns. The composite benchmarks-driven analytical protocol is used as an instrument to examine the understanding, effects, and expert opinions concerning the development patterns and current state of open data ecosystems implemented in eight European countries - Austria, Belgium, Czech Republic, Italy, Latvia, Poland, Serbia, Sweden. 3-round Delphi method is applied to identify, reach a consensus, and validate the observed development patterns and their effects that could lead to disparities and divides. Specifically, this study conducts a comparative analysis of different patterns of open (government) data initiatives and their effects in the eight selected countries using six open data benchmarks, two e-government reports (57 editions in total), and other relevant resources, covering the period of 2013&ndash;2022.</p> <p>***Description of the data in this data set***</p> <p>The file "OpenDataIndex_<em>2013_</em>2022" collects an overview of 27 editions of 6 open data indices - for all countries they cover, providing respective ranks and values for these countries.&nbsp;These indices are:</p> <p>1) Global Open Data Index (GODI) (4 editions)</p> <p>2) Open Data Maturity Report (ODMR) (8 editions)</p> <p>3) Open Data Inventory (ODIN) (6 editions)</p> <p>4) Open Data Barometer (ODB) (5 editions)</p> <p>5) Open, Useful and Re-usable data (OURdata) Index (3 editions)</p> <p>6) Open Government Development Index (OGDI) (2 editions)</p> <p>These data shapes the third context - open data indices and rankings. The second sheet of this file covers countries covered by this study, namely, Austria, Belgium, Czech Republic, Italy, Latvia, Poland, Serbia, Sweden. It serves the basis for Section 4.2 of the paper.</p> <p>Based on the analysis of selected countries, incl. the analysis of their specifics and performance over the years in the indices and benchmarks, covering 57 editions of OGD-oriented reports and indices and e-government-related reports (2013-2022) that shaped a protocol (see paper, Annex 1), 102 patterns that may lead to disparities and divides in the development and benchmarking of ODEs were identified, which after the assessment by expert panel were reduced to a final number of 94 patterns representing four contexts, from which the recommendations defined in the paper were obtained. These patterns are available in the file "OGDdevelopmentPatterns".&nbsp;The first sheet contains the list of patterns, while the second sheet - the list of patterns and their effect as assessed by expert panel.</p> <p>***Format of the file***<br>.xls, .csv (for the first spreadsheet only)</p> <p>***Licenses or restrictions***<br>CC-BY</p> <p>&nbsp;</p> <p>For more info, see README.txt<br>&nbsp;</p>

opencc-by-4.0Nov 2023View details β†’
zenodo44/100

Simulated query reads used for benchmarks in Metabuli publication.

<p>Simulated query reads used for synthetic&nbsp;benchmarks in Metabuli publication.</p><p><strong>Prokaryote Subspecies Inclusion Test</strong></p><ul><li>Illumina paired-end short reads<ul><li>Simulated using Mason2 based on NovaSeq 6000's error rates.</li><li><strong>prokaryote_inclusion_reads_1.fna.gz</strong></li><li><strong>prokaryote_inclusion_reads_2.fna.gz</strong></li></ul></li><li>PacBio HiFi long reads (Sequel II CCS)<ul><li>Multi-pass reads were simulated using PBSIM3 based on Sequel Il's error profile.</li><li>CCS reads were generated using PacBio's CCS-read tool.</li><li><strong>prokaryote_inclusion_sequel_ccs.fq.gz</strong></li></ul></li><li>PacBio Sequel II long reads<ul><li>Simulated using PBSIM3 based on Sequel Il's error profile.</li><li><strong>prokaryote_inclusion_sequel.fq.gz</strong></li></ul></li><li>ONT long reads<ul><li>Simulated using PBSIM3 based on ONT's error profile.</li><li><strong>prokaryote_inclusion_ont.fq.gz</strong></li></ul></li></ul><p><strong>Prokaryote Subspecies exclusion Test</strong></p><ul><li>Illumina paired-end short reads<ul><li>Simulated using Mason2 based on NovaSeq 6000's error rates.</li><li><strong>prokaryote_ss-exclusion_reads_1.fna.gz</strong></li><li><strong>prokaryote_ss-exclusion_reads_2.fna.gz</strong></li></ul></li><li>PacBio HiFi long reads (Sequel II CCS)<ul><li>Multi-pass reads were simulated using PBSIM3 based on Sequel Il's error profile.</li><li>CCS reads were generated using PacBio's CCS-read tool.</li><li><strong>prokaryote_ss-exclusion_sequel_ccs.fq.gz</strong></li></ul></li><li>PacBio Sequel II long reads<ul><li>Simulated using PBSIM3 based on Sequel Il's error profile.</li><li><strong>prokaryote_ss-exclusion_sequel.fq.gz</strong></li></ul></li><li>ONT long reads<ul><li>Simulated using PBSIM3 based on ONT's error profile.</li><li><strong>prokaryote_ss-exclusion_ont.fq.gz</strong></li></ul></li></ul><p><strong>Prokaryote Species Exclusion Test</strong></p><ul><li>Illumina paired-end short reads<ul><li>Simulated using Mason2 based on NovaSeq 6000's error rates.</li><li><strong>prokaryote_exclusion_reads_1.fna.gz</strong></li><li><strong>prokaryote_exclusion_reads_2.fna.gz</strong></li></ul></li><li>PacBio HiFi long reads (Sequel II CCS)<ul><li>Multi-pass reads were simulated using PBSIM3 based on Sequel Il's error profile.</li><li>CCS reads were generated using PacBio's CCS-read tool.</li><li><strong>prokaryote_sp-exclusion_sequel_ccs.fq.gz</strong></li></ul></li><li>PacBio Sequel II long reads<ul><li>Simulated using PBSIM3 based on Sequel Il's error profile.</li><li><strong>prokaryote_sp-exclusion_sequel.fq.gz</strong></li></ul></li><li>ONT long reads<ul><li>Simulated using PBSIM3 based on ONT's error profile.</li><li><strong>prokaryote_exclusion_ont.fq.gz</strong></li></ul></li></ul>

opencc-by-4.0Jun 2023View details β†’
zenodo44/100

VERITE Benchmark

<p>VERITE is a benchmark dataset designed for evaluating multimodal misinformation detection models. The dataset consists of real-world instances of misinformation collected from Snopes and Reuters and addresses unimodal bias by excluding asymmetric misinformation and employing modality balancing.&nbsp;The images are sourced from within the articles of Snopes and Reuters, as well as Google Images. As we do not own the rights to the images, the dataset provide the image URLs along with their captions and labels. VERITE supports multiclass classification of three categories: Truthful, Out-of-context, and Miscaptioned image-caption pairs but can also be used for binary classification. We collected 260 articles from Snopes and 78 from Reuters that met our criteria which translates to 338 Truthful, 338 Miscaptioned and 324 Out-of-Context pairs.&nbsp;</p> <p>For more information on how to use the dataset visit:&nbsp;<a href="https://github.com/stevejpapad/image-text-verification">https://github.com/stevejpapad/image-text-verification</a>.&nbsp;If you encounter any problems while downloading and preparing VERITE (e.g., broken image URLs), please contact&nbsp;<a href="mailto:stefpapad@iti.gr">stefpapad@iti.gr</a>.</p> <p>The dataset was developed in the context of the vera.ai (VERification Assisted by Artificial Intelligence) project.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2024View details β†’
zenodo44/100

CaliParticles: A Benchmark Standard for Experiments in Granular Materials

<p>Granular materials are discrete particulate media that can flow like a liquid but also be rigid like a solid. This complex mechanical behavior originates in part from the particles shape. How particle shape affects mechanical behavior remains poorly understood. Understanding this micro-macro link would enable the rational design of potentially cheap, light weight or robust materials. To aid this development, we have produced a set of standard particle shapes that can be used as benchmarks for granular materials research. Here we describe the collection of benchmark shapes. Some part of the particles are modeled on superquadrics, others are custom designed. The particles used so far were made from polyoxymethylene (POM) and Thermoplastic elastomers (TPE) whose specifications are also listed. The benchmark shapes are available as molds in a plastics manufacturing company, whose contact information is also included. The company is capable of making other molds as well, giving access to more particle shapes. The same particle shapes can thus also be made in different types of (colored) plastic, and in amounts of 50.000 particles or more, larger than conveniently be produced with a 3D printer. We also provide the associated .step and .stl files in the repository in which this document is included.&nbsp;</p>

opencc-by-4.0Oct 2022View details β†’
zenodo44/100

Datasets for benchmarking and ML modelling

<p><em><span>hydrogen-harm</span></em><span> data set of crystalline hydrogen configurations: energies at VMC and LRDMS level; purpose: benchmark for MLP; developed in the group of Michele Casula (CNRS) </span></p> <p><span><em>prot-hex</em> data set for protonated water hexamer: trajectories from classical molecular dynamics with nuclear forces at VMC level of theory; purpose: ML modelling; developed in the group of Michele Casula (CNRS)</span></p> <p><span><em>intexcit</em> data sets for a set of organic molecular complexes in lowest excited states: dispersion interaction energies, interaction energies, components of SAPT interaction energies at the CAS wavefunction level; purpose: benchmarking <em>ab initio</em> methods and density functional dispersion correction modelling; developed by Kasia Pernal (TUL) and Michal Hapka (University of Warsaw) </span></p>

opencc-by-4.0Jan 2024View details β†’
zenodo44/100

Benchmark for classifying presence of coral and camera motion in underwater

<p>Benchmark for classifying presence of coral in underwater videos, and camera motion that would be necessary for 3d reconstruction of coral. Videos are collected from the YouTube-8M dataset.</p>

opencc-by-4.0Mar 2024View details β†’
zenodo44/100

Statistical Process Control Benchmark Dataset

<p>Datasets to the planned publication "Generalized Statistical Process Control via 1D-ResNet Pretraining" by Tobias Schulze, Louis Huebser, Sebastian Beckschulte and Robert H. Schmitt (Chair for Intelligence in Quality Sensing, Laboratory for Machine Tools and Production Engineering, WZL of RWTH Aachen University)</p> <p>Data for benchmarking SPC against other process monitoring methods. The data consist of a one-dimensional timeseries of floats (x.csv). Addititionally information whether the data are within the specifications are provided as another time series (y.csv). The data are generated by solving an optimization problem for each time to generate a mixture distribution of different probability distributions. Then for each timestep one record is sampled. Inputs for the optimization problem are the given probability distributions, the lower and upper limit of the tolerance interval as well as the desired median of the data. Additionally weights of the different probability distributions can be given as boundary condions for the different time steps. Metadata generated from the solving are stored in k_matrix.csv (wheights at each time step) and distribs (probability distribution objects according to https://doi.org/10.5281/zenodo.8249487). The data consists of phases with data from a stable mixture distribution and phases with data from a mixture distribution that do not fulfill the stability criteria.</p> <p>The train data were used to train the G-SPC model. The test data were used for benchmarking purposes</p> <p>Funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany&rsquo;s Excellence Strategy &ndash; EXC-2023 Internet of Production &ndash; 390621612.</p>

openmit-licenseAug 2023View details β†’
zenodo44/100

Results of KROWN: Knowledge Graph Construction Benchmark

<p>In this Zenodo repository we present the results of using KROWN to benchmark popular RDF Graph Materialization systems such as RMLMapper, RMLStreamer, Morph-KGC, SDM-RDFizer, and Ontop (in materialization mode).&nbsp;</p> <h1>What is KROWN πŸ‘‘?</h1> <p>KROWN πŸ‘‘ is a benchmark for materialization systems to construct Knowledge Graphs from (semi-)heterogeneous data sources using declarative mappings such as<a href="http://w3id.org/rml/portal"> RML</a>.</p> <p>Many benchmarks already exist for virtualization systems e.g.<a href="https://github.com/oeg-upm/gtfs-bench"> GTFS-Madrid-Bench</a>,<a href="https://ontop-vkg.org/npd-benchmark/"> NPD</a>,<a href="http://wbsg.informatik.uni-mannheim.de/bizer/berlinsparqlbenchmark/"> BSBM</a> which focus on complex queries with a single declarative mapping. However, materialization systems are unaffected by complex queries since their input is the dataset and the mappings to generate a Knowledge Graph. Some specialized datasets exist to benchmark specific limitations of materialization systems such as duplicated or empty values in datasets e.g.<a href="https://doi.org/10.57702/4c9ivpgs"> GENOMICS</a>, but they do not cover all aspects of materialization systems. Therefore, it is hard to compare materialization systems among each other in general which is where KROWN πŸ‘‘ comes in!&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p> <h1>Results</h1> <p>The raw results are available as ZIP archives, the analysis of the results are available in the spreadsheet <em>results.ods</em>.</p> <h2>Evaluation setup</h2> <p>We generated several scenarios using <a href="https://github.com/kg-construct/KROWN/tree/main/data-generator">KROWN&rsquo;s data generator</a> and executed them 5 times with <a href="https://github.com/kg-construct/KROWN/tree/main/execution-framework">KROWN&rsquo;s execution framework</a>. All experiments were performed on Ubuntu 22.04 LTS machines (Linux 5.15.0, x86_64) with each Intel(R) Xeon(R) CPU E5-2650 v2 @ 2.60GHz, 48 GB RAM memory, and 2 GB swap memory. The output of each materialization system was set to N-Triples.</p> <h2>Materialization systems</h2> <p>We selected the most popular maintained materialization systems for constructing RDF graphs for performing our experiments with KROWN:</p> <ul> <li> <p>RMLMapper</p> </li> <li> <p>RMLStreamer</p> </li> <li> <p>Morph-KGC</p> </li> <li> <p>SDM-RDFizer</p> </li> <li> <p>OntopM (Ontop in materialization mode)</p> </li> </ul> <p><strong>Note</strong>: KROWN is flexible and allows adding any other materialization system, see <a href="https://github.com/kg-construct/KROWN/tree/main/execution-framework">KROWN&rsquo;s execution framework</a> documentation for more information.</p> <h2>Scenarios</h2> <p>We consider the following scenarios:</p> <ul> <li> <p>Raw data: number of rows, columns and cell size</p> </li> <li> <p>Duplicates &amp; empty values: percentage of the data containing duplicates or empty values</p> </li> <li> <p>Mappings: Triples Maps (TM), Predicate Object Maps (POM), Named Graph Maps (NG).</p> </li> <li> <p>Joins: relations (1-N, N-1, N-M), conditions, and duplicates during joins</p> </li> </ul> <p><strong>Note</strong>: KROWN is flexible and allows adding any other scenario, see <a href="https://github.com/kg-construct/KROWN/tree/main/data-generator">KROWN&rsquo;s data generator documentation</a> for more information.</p> <p>In the table below we list all parameter values we used to configure our scenarios:</p> <div> <table> <tbody> <tr> <td> <p><strong>Scenario</strong></p> </td> <td> <p><strong>Parameter values</strong></p> </td> </tr> <tr> <td> <p>Raw data: rows</p> </td> <td> <p>10K, 100K, 1M, 10M</p> </td> </tr> <tr> <td> <p>Raw data: columns</p> </td> <td> <p>1, 10, 20, 30</p> </td> </tr> <tr> <td> <p>Raw data: cell size</p> </td> <td> <p>500, 1K, 5K, 10K&nbsp;</p> </td> </tr> <tr> <td> <p>Duplicates: percentage</p> </td> <td> <p>0%, 25%, 50%, 75%, 100%</p> </td> </tr> <tr> <td> <p>Empty values: percentage</p> </td> <td> <p>0%, 25%, 50%, 75%, 100%</p> </td> </tr> <tr> <td> <p>Mappings: TMs + 5POMs</p> </td> <td> <p>1, 10, 20, 30 TMs</p> </td> </tr> <tr> <td> <p>Mappings: 20TMs + POMs</p> </td> <td> <p>1, 3, 5, 10 POMs</p> </td> </tr> <tr> <td> <p>Mappings: NG in SM</p> </td> <td> <p>1, 5, 10, 15 NGs</p> </td> </tr> <tr> <td> <p>Mappings: NG in POM</p> </td> <td> <p>1, 5, 10, 15 NGs</p> </td> </tr> <tr> <td> <p>Mappings: NG in SM/POM</p> </td> <td> <p>1/1, 5/5, 10/10, 15/15 NGs</p> </td> </tr> <tr> <td> <p>Joins: 1-N relations</p> </td> <td> <p>1-1, 1-5, 1-10, 1-15</p> </td> </tr> <tr> <td> <p>Joins: N-1 relations</p> </td> <td> <p>1-1, 5-1, 10-1, 15-1</p> </td> </tr> <tr> <td> <p>Joins: N-M relations&nbsp;</p> </td> <td> <p>3-3, 3-5, 5-3, 10-5, 5-10</p> </td> </tr> <tr> <td> <p>Joins: join conditions</p> </td> <td> <p>1, 5, 10, 15</p> </td> </tr> <tr> <td> <p>Joins: join duplicates</p> </td> <td> <p>0, 5, 10, 15</p> </td> </tr> </tbody> </table> </div> <h1>&nbsp;</h1>

opencc-by-4.0Apr 2024View details β†’
zenodo44/100

Simulated BCRseq data for the nf-core/airrflow pipeline benchmark

<p>Simulated BCRseq data for the nf-core/airrflow pipeline benchmark.</p> <p>The original repertoires simulated with ImmuneSim (sim_repertoire_orig_repA/repB/repC.tsv) and the clonally expanded repertoires (sim_repertoire_clonally_expandedrepA/repB/repC.tsv), as well as the fasta file formats from the clonally expanded repertoires with and without UMIs are also shared. The fasta files were used to simulate the sequencing reads with various degrees of sequencing errors.</p>

opencc-by-4.0May 2023View details β†’
zenodo44/100

Bio-logger Ethogram Benchmark: A benchmark for computational analysis of animal behavior, using animal-borne tags

<p>This repository contains the datasets and experiment results presented in our <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a>:</p> <blockquote> <p>B. Hoffman, M. Cusimano, V. Baglione, D. Canestrari, D. Chevallier, D. DeSantis, L. Jeantet, M. Ladds, T. Maekawa, V. Mata-Silva, V. Moreno-Gonz&aacute;lez, A. Pagano, E. Trapote, O. Vainio, A. Vehkaoja, K. Yoda, K. Zacarian, A. Friedlaender, "A benchmark for computational analysis of animal behavior, using animal-borne tags," 2023.</p> </blockquote> <p>Standardized code to implement, train, and evaluate models can be found at <a href="https://github.com/earthspecies/BEBE/">https://github.com/earthspecies/BEBE/</a>.&nbsp;</p> <p>Please note the licenses in each dataset folder.</p> <p><strong>Zip folders beginning with "formatted":</strong> These are the datasets we used to run the experiments reported in the benchmark paper.&nbsp;</p> <p><strong>Zip folders beginning with "raw": </strong>These are the unprocessed datasets used in BEBE. Code to process&nbsp;these raw datasets into the formatted ones used by BEBE can be found at&nbsp;<a href="https://github.com/earthspecies/BEBE-datasets/">https://github.com/earthspecies/BEBE-datasets/</a>.</p> <p><strong>Zip folders beginning with "experiments": </strong>Results of the cross-validation experiments reported in the paper, as well as hyperparameter optimization. Confusion matrices for all experiments can also be found here. Note that dt, rf, and svm refer to the feature set from Nathan et al., 2012.</p> <p><em>Results used in Fig. 4 of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a> (deep neural networks vs. classical models)</em><br>{dataset}_ harnet_nogyr<br>{dataset}_CRNN<br>{dataset}_CNN<br>{dataset}_dt<br>{dataset}_rf<br>{dataset}_svm<br>{dataset}_wavelet_dt<br>{dataset}_wavelet_rf<br>{dataset}_wavelet_svm</p> <p><em>Results used in Fig. 5D of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a> (full data setting)<br></em>If dataset contains gyroscope (HAR, jeantet_turtles, vehkaoja_dogs):<br>{dataset}_harnet_nogyr<br>{dataset}_harnet_random_nogyr<br>{dataset}_harnet_unfrozen_nogyr<br>{dataset}_RNN_nogyr<br>{dataset}_CRNN_nogyr<br>{dataset}_rf_nogyr<br><br>Otherwise:<br>{dataset}_harnet_nogyr<br>{dataset}_harnet_unfrozen_nogyr<br>{dataset}_harnet_random_nogyr<br>{dataset}_RNN_nogyr<br>{dataset}_CRNN<br>{dataset}_rf</p> <p><em>Results used in Fig. 5E of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a> (reduced data setting)<br></em>If dataset contains gyroscope (HAR, jeantet_turtles, vehkaoja_dogs):<br>{dataset}_harnet_low_data_nogyr<br>{dataset}_harnet_random_low_data_nogyr<br>{dataset}_harnet_unfrozen_low_data_nogyr<br>{dataset}_RNN_low_data_nogyr<br>{dataset}_wavelet_RNN_low_data_nogyr<br>{dataset}_CRNN_low_data_nogyr<br>{dataset}_rf_low_data_nogyr</p> <p>Otherwise:<br>{dataset}_harnet_low_data_nogyr<br>{dataset}_harnet_random_low_data_nogyr<br>{dataset}_harnet_unfrozen_low_data_nogyr<br>{dataset}_RNN_low_data_nogyr<br>{dataset}_wavelet_RNN_low_data_nogyr<br>{dataset}_CRNN_low_data<br>{dataset}_rf_low_data<br><br></p> <p><strong>CSV files</strong>: we also include summaries of the experimental results in experiments_summary.csv, experiments_by_fold_individual.csv, experiments_by_fold_behavior.csv.&nbsp;</p> <p><em>experiments_summary.csv - results averaged over individuals and behavior classes<br></em>dataset (str): name of dataset<br>experiment (str): name of model with experiment setting&nbsp;<br>fig4 (bool): True if dataset+experiment was used in figure 4 of&nbsp;<a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5d (bool): True if dataset+experiment was used in figure 5d of&nbsp;<a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5e (bool): True if dataset+experiment was used in figure 5e of&nbsp;<a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>f1_mean (float): mean of macro-averaged F1 score, averaged over individuals in test folds<br>f1_std (float): standard deviation of macro-averaged F1 score, computed over individuals in test folds<br>prec_mean, prec_std (float): analogous for precision<br>rec_mean, rec_std (float): analogous for recall<em><br><br>experiments_by_fold_individual.csv - results per individual in the test folds<br></em>dataset (str): name of dataset<br>experiment (str): name of model with experiment setting&nbsp;<br>fig4 (bool): True if dataset+experiment was used in figure 4 of&nbsp;<a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5d (bool): True if dataset+experiment was used in figure 5d of&nbsp;<a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5e (bool): True if dataset+experiment was used in figure 5e of&nbsp;<a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fold (int): test fold index<br>individual (int): individuals are numbered zero-indexed, starting from fold 1<br>f1 (float): macro-averaged f1 score for this individual<br>precision (float): macro-averaged precision for this individual<br>recall (float): macro-averaged recall for this individual<em><br></em></p> <p><em>experiments_by_fold_behavior.csv - results per behavior class, for each test fold<br></em>dataset (str): name of dataset<br>experiment (str): name of model with experiment setting&nbsp;<br>fig4 (bool): True if dataset+experiment was used in figure 4 of&nbsp;<a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5d (bool): True if dataset+experiment was used in figure 5d of&nbsp;<a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5e (bool): True if dataset+experiment was used in figure 5e of&nbsp;<a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fold (int): test fold index<br>behavior_class (str): name of behavior class<br>f1 (float): f1 score for this behavior, averaged over individuals in the test fold<br>precision (float): precision for this behavior, averaged over individuals in the test fold<br>recall (float): recall for this behavior, averaged over individuals in the test fold<br>train_ground_truth_label_counts (int): number of timepoints labeled with this behavior class, in the training set<em><br></em></p>

opencc-by-4.0Apr 2024View details β†’
zenodo44/100

SANTOS Benchmark for Table Union Search

<p>This record contains the datasets released with&nbsp;<a href="https://2023.sigmod.org/">SIGMOD 2023</a> paper entitled "<a href="https://dl.acm.org/doi/10.1145/3588689">SANTOS: Relationship-based Semantic Table Union Search</a>". We release two new tabular&nbsp;benchmarks to evaluate the table union search problem&nbsp;over the data lakes. Furthermore, we also release relabeled ground truth for an existing <a href="https://github.com/RJMillerLab/table-union-search-benchmark">TUS benchmark</a> by taking the binary relationship between the columns into account. Please visit <a href="https://dl.acm.org/doi/10.1145/3588689">our paper</a> for further details.</p> <p>If you use our dataset for your work, please cite our paper as:</p> <p>Aamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen, Wolfgang Gatterbauer, Ren&eacute;e J. Miller, and Mirek<br>Riedewald. 2023. SANTOS: Relationship-based Semantic Table Union Search. SIGMOD Conference 2023, ACM</p> <p>@article{DBLP:journals/pacmmod/KhatiwadaFSCGMR23,<br>&nbsp; author &nbsp; &nbsp; &nbsp; = {Aamod Khatiwada and<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Grace Fan and<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Roee Shraga and<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Zixuan Chen and<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Wolfgang Gatterbauer and<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Ren{\'{e}}e J. Miller and<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Mirek Riedewald},<br>&nbsp; title &nbsp; &nbsp; &nbsp; &nbsp;= {{SANTOS:} Relationship-based Semantic Table Union Search},<br>&nbsp; journal &nbsp; &nbsp; &nbsp;= {Proc. {ACM} Manag. Data},<br>&nbsp; volume &nbsp; &nbsp; &nbsp; = {1},<br>&nbsp; number &nbsp; &nbsp; &nbsp; = {1},<br>&nbsp; pages &nbsp; &nbsp; &nbsp; &nbsp;= {9:1--9:25},<br>&nbsp; year &nbsp; &nbsp; &nbsp; &nbsp; = {2023},<br>&nbsp; doi &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;= {10.1145/3588689},<br>}</p> <p>You can find SANTOS implementation at:&nbsp;<a href="https://github.com/northeastern-datalab/santos">https://github.com/northeastern-datalab/santos</a></p> <p>You can find the original TUS benchmark at:&nbsp;<a href="https://github.com/RJMillerLab/table-union-search-benchmark">https://github.com/RJMillerLab/table-union-search-benchmark</a></p> <p>Abstract:&nbsp;Existing techniques for unionable table search define unionability using metadata (tables must have the same or similar schemas) or column-based metrics (for example, the values in a table should be drawn from the same domain). In this work, we introduce the use of semantic relationships between pairs of columns in a table to improve the accuracy of union search. Consequently, we introduce a new notion of unionability that considers relationships between columns, together with the semantics of columns, in a principled way. To do so, we present two new methods to discover semantic relationship between pairs of columns: The first uses an existing knowledge base (KB), the second (which we call a &ldquo;synthesized KB&rdquo;) uses knowledge from the data lake itself. We adopt an existing Table Union Search benchmark and present new (open) benchmarks that represent small and large real data lakes. We show that our new unionability search algorithm called SANTOS outperforms a state-of-the-art union search that uses a wide variety of column-based semantics, including word embeddings and regular expressions. We show empirically in all benchmarks that our synthesized KB improves the accuracy of union search by representing relationship semantics that may not be contained in an available KB. This result hints at a promising future of creating a synthesized KBs from data lakes with limited KB coverage and using them for union search.</p>

opencc-by-4.0Mar 2023View details β†’
zenodo44/100

On CNF Conversion for SAT and SMT Enumeration: Benchmarks, Results and Plots

<p>Experimental results for the paper:<br><br><a title="Arxiv Link" href="https://arxiv.org/abs/2303.14971" target="_blank" rel="noopener">On CNF encoding for SAT and SMT enumeration</a>, Gabriele Masina, Giuseppe Spallitta and Roberto Sebastiani. ArXiv, 2024.</p> <p>Content:</p> <ul> <li><code>aig-bench.zip</code> <code>iscas85-bench.zip</code> <code>syn-bool-bench.zip</code> contain the inputs and results for the Boolean benchmarks. Each zip contains: <ul> <li><code>data/</code> that contains the input data</li> <li><code>results-&lt;tool&gt;/</code> for each tested tool.</li> </ul> </li> <li><code>syn-lra-bench.zip</code> <code>wmi-bench.zip</code> contain the inputs and results for the Boolean benchmarks. Each zip contains: <ul> <li><code>data/</code> that contains the input data</li> <li><code>results-&lt;tool&gt;/</code> for each tested tool.</li> </ul> </li> <li><code>plot-d4</code>,&nbsp;<code>plot-msat</code>,&nbsp;<code>plot-tabularallsat</code>,&nbsp;<code>plot-tabularallsmt</code> contain the plots for enumeration with different tools.&nbsp;</li> <li><code>plot-msat-sat</code> contains the plot for plain satisfiability using MathSAT.</li> </ul> <p>Results are stored in JSON files, where the field <code>"mode"</code> indicates the CNF transformation used to preprocess the input:</p> <ul> <li><code>LAB</code> for Tseitin CNF</li> <li><code>LABELNEG_POL</code> for Plaisted&amp;Greenbaum CNF, using negative labels for subformulas occurring negatively only.</li> <li><code>NNF_MUTEX_POL</code> for NNF+Plaisted&amp;Greenbaum CNF+mutex clauses, as described in the paper</li> </ul> <p>The source code used to run the experiments is available at&nbsp;<a href="https://doi.org/10.5281/zenodo.14033422" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.14033422</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details β†’
zenodo44/100

A data set from an extensive experimental benchmark study of the Hell Bridge Test Arena subject to imposed damage

<p>A data set from an extensive experimental benchmark study of the Hell Bridge Test Arena (HBTA), a full-scale steel bridge subject to imposed damage, has been established. The data set includes organized dynamic response and load measurement data of the bridge under different structural state conditions, where the structural state conditions range from an undamaged (reference) state to known damage states. Furthermore, the data set includes acceleration and strain data from the response monitoring and acceleration data from the load monitoring, where a modal vibration shaker is used as an excitation source. The data is collected in one h5-file (hierarchical data format version 5) with a sampling rate of 100 Hz. Signal processing and resampling of the data has been performed according to the description provided in the references below. The data set is now published in this open-access data repository and can be accessed and downloaded freely. As such, the data set provides an important benchmark to the scientific community within bridge damage detection and SHM.</p>

opencc-by-4.0Jan 2024View details β†’
zenodo44/100

L'Alpe d'Huez: a dataset to benchmark topographic map generalisation

<p>This dataset derives from the one used in a past EuroSDR benchmark (http://dx.doi.org/10.1016/j.compenvurbsys.2009.06.002), and should be used as a benchmark for topographic map generalisation techniques. It contains several topographic layers as shapefiles (roads, buildings, rivers, forests, contour lines...), a style description to display the data at the 1:50k scale, and a table of the generalisation constraints that should be respected in this 1:50k scale map.</p> <p>The initial data can considered as detailed for maps at the 1:15k scale. The projection of the data is "Lambert II Etendu", EPSG:27572.</p> <p>The area is 11*11 km large.</p>

opencc-by-4.0Sep 2021View details β†’

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record