Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
200
datasets available to search
ShareScore release 0.9.0
Dataset results
200 results for “accuracy evaluation”
An evaluation of the accuracy and speed of metagenome analysis tools
<p>Metagenome studies are becoming increasingly widespread, yielding important insights into microbial communities covering diverse environments from terrestrial and aquatic ecosystems to human skin and gut. With the advent of high-throughput sequencing platforms, the use of large scale shotgun sequencing approaches is now commonplace. However, a thorough independent benchmark comparing state-of-the-art metagenome analysis tools is lacking. Here, we present a benchmark where the most widely used tools are tested on complex, realistic data sets. Our results clearly show that the most widely used tools are not necessarily the most accurate, that the most accurate tool is not necessarily the most time consuming and that there is a high degree of variability between available tools. These findings are important as the conclusions of any metagenomics study are affected by errors in the predicted community composition and functional capacity.</p>
Simulation data for eddy current rail testing - simulation accuracy and evaluation uncertainty quantification
<p>This dataset serve to quantify the simulation error and the evaluation uncertainties in the context of eddy current rail testing. It was obtained during the AIFRI project (Artificial Intelligence for Rail Inspection) with the Faraday software by INTEGRATED Engineering Software, using its BEM Solver. For an analysis see the article below.</p>
Geomorpho90m, empirical evaluation and accuracy assessment of global high-resolution geomorphometric layers
<p>Topographical relief comprises the vertical and horizontal variations of the Earth’s terrain and drives processes in geomorphology, biogeography, climatology, hydrology and ecology. Its characterisation and assessment, through geomorphometry and feature extraction, is fundamental to numerous environmental modelling and simulation analyses. We, therefore, developed the Geomorpho90m global dataset comprising of different geomorphometric features derived from the MERIT-Digital Elevation Model (DEM) - the best global, high-resolution DEM available. The fully-standardised 26 geomorphometric variables consist of layers that describe the (i) rate of change across the elevation gradient, using first and second derivatives, (ii) ruggedness, and (iii) geomorphological forms. The Geomorpho90m variables are available at 3 (~90 m) and 7.5 arc-second (~250 m) resolutions under the WGS84 geodetic datum, and 100 m spatial resolution under the Equi7 projection. They are useful for modelling applications in fields such as geomorphology, geology, hydrology, ecology and biogeography.</p> <p>Publication <a href="https://www.nature.com/articles/s41597-020-0479-6">https://www.nature.com/articles/s41597-020-0479-6</a></p>
Geomorpho90m, empirical evaluation and accuracy assessment of global high-resolution geomorphometric layers
<p>Topographical relief comprises the vertical and horizontal variations of the Earth’s terrain and drives processes in geomorphology, biogeography, climatology, hydrology and ecology. Its characterisation and assessment, through geomorphometry and feature extraction, is fundamental to numerous environmental modelling and simulation analyses. We, therefore, developed the Geomorpho90m global dataset comprising of different geomorphometric features derived from the MERIT-Digital Elevation Model (DEM) - the best global, high-resolution DEM available. The fully-standardised 26 geomorphometric variables consist of layers that describe the (i) rate of change across the elevation gradient, using first and second derivatives, (ii) ruggedness, and (iii) geomorphological forms. The Geomorpho90m variables are available at 3 (~90 m) and 7.5 arc-second (~250 m) resolutions under the WGS84 geodetic datum, and 100 m spatial resolution under the Equi7 projection. They are useful for modelling applications in fields such as geomorphology, geology, hydrology, ecology and biogeography.</p> <p>Publication <a href="https://www.nature.com/articles/s41597-020-0479-6">https://www.nature.com/articles/s41597-020-0479-6</a></p>
Recordings from: Evaluation of a coastal acoustic buoy for cetacean detections, bearing accuracy, and exclusion zone monitoring
<p>1.<span> </span>There is strong socio-political support for offshore wind development in US territorial waters, and construction is planned off several east coast states. Some of the planned development sites coincide with important habitat for critically endangered North Atlantic right whales. Both exclusion zones and passive acoustic monitoring are important tools for managing interactions between marine mammals and human activities. Understanding where animals are with respect to exclusion zones is important to avoid costly construction delays while minimizing the potential for negative impacts. Impact piling from construction of hundreds of offshore wind turbines likely requires exclusion zones as large as 10 km.</p> <p>2.<span> </span>We have developed a three-hydrophone passive acoustic monitoring system that provides bearing information along with marine mammal detections to allow for informed management decisions in real-time. Multiple units form a monitoring system designed to determine whether marine mammal calls originate from inside or outside of an exclusion zone. In October 2021 we undertook a full system validation, with a focus on evaluating the detection range and bearing accuracy of the system with respect to right whale upcalls. Five units were deployed in Mid-Atlantic waters and we played more than >3,500 simulated right whale upcalls at known locations to characterize the detection function and bearing accuracy of each unit. The modeled results of the detection function error were then used to compare the effectiveness of a bearing-based system to a single sensor that can only detect a signal but not ascertain directivity.</p> <p>3.<span> </span>Field trials indicated maximum detection ranges from 4–7.3 km depending on source and ambient noise levels. Simulations showed that incorporating bearing detections provides a substantial improvement in false alarm rates (6 to 12 times depending on number of units, placement, and signal to noise conditions) for a small increase in the risk of missed detections inside of an exclusion zone (1–3%). </p> <p>4.<span> </span>We show that the system can be used for monitoring exclusion zones and clearly highlight the value of including bearing estimation into exclusion zone monitoring plans while noting that placement and configuration of units should reflect anticipated ambient noise conditions.</p>
Datasets to Evaluate Accuracy, Miscalibration and Popularity Lift in Recommendations
<p>This repository contains three datasets for evaluating accuracy, miscalibration and popularity lift in recommender systems. All datasets contain genre/category information in addition to different user group splits:</p> <ol> <li>Last.fm (lfm.zip), based on the LFM-1b dataset of JKU Linz (http://www.cp.jku.at/datasets/LFM-1b/)</li> <li>MovieLens (ml.zip), based on MovieLens-1M dataset (https://grouplens.org/datasets/movielens/1m/)</li> <li>MyAnimeList (anime.zip), based on the MyAnimeList dataset of Kaggle (https://www.kaggle.com/CooperUnion/anime-recommendations-database)</li> </ol> <p>'user_events_cats.txt' contains the users' rating/interaction data along with a list of genres/categories assigend to the rated items. The list of categories is given in 'categories.txt'. Additionally, assignments to three user groups that differ in their inclination to popular/mainstream items are provided: LowPop in 'low_main_users.txt', MedPop in 'med_main_users.txt', and HighPop in 'high_main_users.txt'.</p> <p>The format of the three user files are "user,mainstreaminess"</p> <p>The format of the user-events files are "user,item,preference,cats", where different categories are separated by '|'</p> <p>The format of the categories files are "category-name,index", where index refers to the category-id in the user-events files</p> <p>Example Python-code for analyzing the datasets as well as empirical results on calibration, popularity lift and accuracy can be found on GitHub: https://github.com/domkowald/FairRecSys</p>
Dataset for the publication "Evaluation of Voltage Transformers' Accuracy in Harmonic and Interharmonic Measurement"
<p>This is dataset for paper published:</p> <p>G. Crotti, G. D’Avanzo, C. Landi, P. S. Letizia and M. Luiso, "Evaluation of Voltage Transformers’ Accuracy in Harmonic and Interharmonic Measurement," in <em>IEEE Open Journal of Instrumentation and Measurement</em>, vol. 1, pp. 1-10, 2022, Art no. 9000310, doi: 10.1109/OJIM.2022.3198473.</p>
BQE WIM Data Year 6 Project (Evaluation of Integrated Overweight Enforcement System using High Accuracy WIM System and Non-Proprietary ALPR System)
<p>NEW BQE (Brooklyn-Queens Expressway) WIM Data for QB (Queens Bound) for Direct Overweight Enforcement</p>
Isotope mixing scenarios for: To what extent are the source mixing models accurate: evaluation of the model accuracy and guidelines for the site-specific model selection
<p><span>We selected 10 types of distinct isotope signatures that can be found in the samples of natural water. Every 3–10 types of hypothetical isotope signatures were conceptually grouped together. There would be 968 possible combinations based on combinatorics theory. However, we needed distinct mixing polygons to facilitate our determination of model capacity in dealing with uncertainties. Therefore, we </span><span>kept </span><span>only 240 such groups in </span><span>the </span><span>final</span><span> analysis</span><span>. Each group was designated with a </span><span>predefined</span><span> mixing ratio. After that, we ran all the examined models through these mixing scenarios to </span><span>obtain</span><span> the model estimation of the mixing ratios.</span></p>
Recordings from: Evaluation of a coastal acoustic buoy for cetacean detections, bearing accuracy, and exclusion zone monitoring
Open the record for dataset details and reuse information.
Isotope mixing scenarios and machine learning model in: To what extent are the source mixing models accurate: evaluation of the model accuracy and guidelines for the site-specific model selection
Open the record for dataset details and reuse information.
An evaluation of the accuracy and speed of metagenome analysis tools
<p>Metagenome studies are becoming increasingly widespread, yielding important insights into microbial communities covering diverse environments from terrestrial and aquatic ecosystems to human skin and gut. With the advent of high-throughput sequencing platforms, the use of large scale shotgun sequencing approaches is now commonplace. However, a thorough independent benchmark comparing state-of-the-art metagenome analysis tools is lacking. Here, we present a benchmark where the most widely used tools are tested on complex, realistic data sets. Our results clearly show that the most widely used tools are not necessarily the most accurate, that the most accurate tool is not necessarily the most time consuming and that there is a high degree of variability between available tools. These findings are important as the conclusions of any metagenomics study are affected by errors in the predicted community composition and functional capacity.</p>
Supplementary Table S1. Combined analysis of variance containing the degrees of freedom (DF), mean squares (MS), P value (P val.), mean, coefficient of experimental variation (CEV%) and selective accuracy (SA) for the traits of luminosity (L*), chromaticity a* (a*), chromaticity b* (b*), grain length (length, mm), grain width (width, mm), grain thickness (thickness, mm), mass of 100 grains (Mass, g), normal grains (Ng, %), water absorption (absorption, %), cooking time (Ct, min:s), and concentrations of potassium (K, g kg-1 dry matter - DM), phosphorus (P, g kg-1 DM), calcium (Ca, g kg-1 DM), magnesium (Mg, g kg-1 DM), iron (Fe, mg kg-1 DM), zinc (Zn, mg kg-1 DM), and copper (Cu, mg kg-1 DM) obtained in 25 common bean cultivars evaluated in four experiments carried out from 2019 to 2021
<p><strong><span>Table S1.</span></strong><span> Combined analysis of variance.</span></p> <p><strong><span>Indirect selection for multiple technological and nutritional traits in common bean cultivars under different degrees of multicollinearity</span></strong></p> <p><strong><span>Bragantia, 2024.</span></strong></p>
Data set for the publication " Evaluation of the Accuracy and Frequency Response of Medium-Voltage Instrument Transformers under the Combined Influence Factors of Temperature and Vibration"
<p>This is dataset for paper published:</p> <p>Agazar, M.; Istrate, D.; Pradayrol, P. Evaluation of the Accuracy and Frequency Response of Medium-Voltage Instrument Transformers under the Combined Influence Factors of Temperature and Vibration. <em>Energies</em> <strong>2023</strong>, <em>16</em>, 5012. https://doi.org/10.3390/en16135012</p>
Balancing Accuracy and Evaluation Overhead in Simulation Point Selection
<p>This is the dataset collected for the paper "Balancing Accuracy and Evaluation Overhead in Simulation Point Selection" published at IISWC 2023. It contains 710850 SimPoint configurations and their statistics collected over 22 SPEC2017 benchmarks.</p>
Data from: Evaluating the accuracy of methods for detecting correlated rates of molecular and morphological evolution
<p class="MsoNormal"><span>Determining the link between genomic and phenotypic change is a </span><span>fundamental goal in evolutionary biology. Insights into this link can be gained by using a phylogenetic approach to test for correlations between rates of molecular and morphological evolution. However, there has been persistent uncertainty about the relationship between these rates, partly because conflicting results have been obtained using various methods that have not been examined in detail. We carried out a simulation study to evaluate the performance of five statistical methods for detecting correlated rates of evolution. Our simulations explored the evolution of molecular sequences and morphological characters under a range of conditions. Of the methods tested, Bayesian relaxed-clock estimation of branch rates was able to detect correlated rates of evolution correctly in the largest number of cases. This was followed by correlations of root-to-tip distances, Bayesian model selection, independent sister-pairs contrasts, and likelihood-based model selection. As expected, the power to detect correlated rates increased with the amount of data, both in terms of tree size and number of morphological characters. Likewise, greater among-lineage rate variation in the data led to improved performance of all five methods, particularly for Bayesian relaxed-clock analysis when the rate model was mismatched. We then applied these methods to a data set from flowering plants and did not find evidence of a correlation in evolutionary rates between genomic data and morphological characters. The results of our study have practical implications for phylogenetic analyses of combined molecular and morphological data sets, and highlight the conditions under which the links between genomic and phenotypic rates of evolution can be evaluated quantitatively.</span></p>
In Vivo Evaluation of the Accuracy of Immediate Screw-Retained Provisional Crowns Fabricated Using Digital Planning and Guided Surgery
ClinicalTrials.gov study NCT07315607. IPD Sharing: NO. Countries: 1. Publications: 19.
Evaluating Accuracy, Impact, and Operational Challenges of GeneXpert Use for TB Case Finding Among HIV-infected Persons
ClinicalTrials.gov study NCT02538952. IPD Sharing: UNDECIDED. Countries: 1. Publications: 4.
Evaluation of the Accuracy of a Computer Vision-based Tool for Assessment of Total Body Fat Percentage
ClinicalTrials.gov study NCT04854421. IPD Sharing: NO. Countries: 1. Publications: 1.
Evaluating the Clinical Accuracy of Gallium-68 PSMA PET/CT Imaging in Patients With Biochemical Recurrence of Prostate Cancer
ClinicalTrials.gov study NCT03822845. IPD Sharing: YES. Countries: 1. Publications: 5.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.