Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,503
datasets available to search
ShareScore release 0.7.1
Dataset results
7,503 results for “methods”
Accompanying dataset for: A Monte Carlo Method for Metamorphic Testing of Machine Translation Services
<p>This dataset includes enhanced analysis of the machine translation data. The original dataset has been reported in [1], where white spaces were used to separate words in different languages. This is however not the best method of analyzing some Asian languages such as the Chinese language. In the present analysis, we used a character-based approach to separating the Chinese and Japanese results, hence obtaining a different set of BLEU and Cosine Similarity scores. These new scores are given in the present dataset.</p> <p>[1] Daniel Pesu, Zhi Quan Zhou, Jingfeng Zhen, & Dave Towey. (2018). Accompanying dataset for: A Monte Carlo Method for Metamorphic Testing of Machine Translation Services (Version 1.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1194560</p>
Method Classification of Open Access INTACT Molecular Interaction data.
<p>Simple classification data derived from open access papers indexed in the INTACT database (https://www.ebi.ac.uk/intact/downloads) based on PSI-MI25 codes for interaction detection methods or participant detection methods based on the subfigure caption text. <br> <br> intact_records_and_captions_complete.tsv - This file links available text of subfigure captions to PSI-MI25 codes for the interaction detection method and participant detection method. </p> <p>evidx_run_file.txt - This file provides execution codes for the 'EvidX' machine learning text classifier (https://github.com/SciKnowEngine/evidX/releases/tag/v0.1.0)</p> <p> </p> <p> </p> <p> </p>
R-code for publication: Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods
<p>This is the R-code as well as the underlying data needed to reproduce the results of the springer book chapter: "Ensembles of Ensembles: Combining the Predictions from Multiple Machine Learning Methods"</p> <p>For more information contact: dlieske@mta.ca</p> <p> </p>
Integrated pedagogical methods effectiveness in Physics' preliminary undergraduate education within the context of large size lectures.
<p>Three files relating the first round of analysis testing active methods for large size lectures. The Presentation including the research design, methods, main results in synthesis is available here: https://www.researchgate.net/project/Getting-started-with-Physics-preliminary-undergraduated-strategies/update/5a44cac6b53d2f0bba475104</p> <p>2- Dataset on Students' Learning Outcomes. Dataset adopted in the first experimental round. The dataset includes data used for the first type of analysis (learning outcomes) carried out for the ICEM2017 Conference presentation "Integrating MOOCs in Physics preliminary undergraduate education: beyond large size lectures". The data includes the results of the initial, baseline Test, the final Test, and two other variables that could be used to analyse covariance: Sex and Type of Group (Large/Small).</p> <p>3- Dataset on Students' Opinion. Dataset adopted in the first experimental round. The dataset includes data used for the second type of analysis (students' opinion) carried out for the ICEM2017 Conference presentation "Integrating MOOCs in Physics preliminary undergraduate education: beyond large size lectures". The data includes the results of a final questionnaire gathering the students opinion on the four types of pedagogical factors affecting their experience within a large size lecture: MOOCs, Active Learning, Self-Assessment tools, Tutors’ guidance.</p> <p>4- Codes and analysis adopted in the first experimental round. The Document includes two analysis carried on for the ICEM2017 Conference presentation "Integrating MOOCs in Physics preliminary undergraduate education: beyond large size lectures". These are: Test (measuring students' knowledge on the subject taught) and Students' Opinion/satisfaction on the several pedagogical methods adopted along the experimental intervention.</p>
Supplementary Material for "Formal methods in dependable systems engineering: a survey of professionals from Europe and North America"
<p>This report contains supplemental material for <a href="https://link.springer.com/article/10.1007%2Fs10664-020-09836-5">this paper</a>, including a detailed analysis of responses to certain questions, further visualizations of the collected data, details on our analysis of related work, and a copy of the whole questionnaire. This material was shared for the period of peer review and has been significantly updated, extended, and included in <a href="https://link.springer.com/article/10.1007%2Fs10664-020-09836-5">this journal publication</a>.</p>
A decade of Semantic Web research through the lenses of a mixed methods approach (Resources)
<p>This work has been submitted to <a href="http://www.semantic-web-journal.net/content/decade-semantic-web-research-through-lenses-mixed-methods-approach">Semantic Web Journal</a>. We provide here resources to reproduce our approach.</p> <p>In this paper, we aim to provide a broader and more complete picture of Semantic Web topics and trends by adopting a mixed methods methodology, which allows a combined use of both qualitative and quantitative approaches. Concretely, we build on a qualitative analysis of the main seminal papers, which adopt a top-down approach, and on quantitative results derived with three bottom-up data-driven approaches (<a href="https://technologies.kmi.open.ac.uk/Rexplore/">Rexplore</a>, <a href="http://saffron.insight-centre.org/">Saffron</a>, <a href="https://www.poolparty.biz/">PoolParty</a>), on a corpus of Semantic Web papers published in the last decade. In this process, we both use the latter for “fact-checking” on the former and also to derive key findings in relation to the strengths and weaknesses of top-down and bottom-up approaches to research topic identification.</p> <p>Please access the full set of resources at: <a href="https://aic.ai.wu.ac.at/qadlod/SW/">https://aic.ai.wu.ac.at/qadlod/SW/</a></p>
Recognising innovative companies by using a diversified stacked generalisation method for website classification – the raw results
<p><strong>Introduction</strong></p> <p>The classification models were trained out by using the Classification and Regression Training package (caret) [1]. The models' parameters were fine-tuned by the 10-fold cross-validation procedure [2].</p> <p><strong>Cluster parameters</strong></p> <p>Most computations were carried out on a cluster having the following parameters:</p> <ul> <li>GPU: NVIDIA Tesla P100;</li> <li>CPU: 2.0 GHz Intel® Xeon® Platinum 8167M;</li> <li>The number of GPUs: 2;</li> <li>The number of CPU cores: 28;</li> <li>The number of CPU threads: 56;</li> <li>RAM: 192 GB;</li> <li>Storage: 3 TB.</li> </ul> <p>Only one model (k-nn) was calculated on a cluster having the following parameters:</p> <ul> <li>Processor: Intel(R) Core(TM) i7-4770 CPU @ 3.40GHz 3.40 GHz;</li> <li>RAM: 16 GB;</li> <li>Windows 64 bit.</li> </ul> <p><strong>Performance statistics</strong></p> <p>All performance statistics are stored in cvs files. Each file corresponds to a particular machine learning method such as a file, "methodName-stat.csv" contains all data regarding a method, "methodName." All files cover the following columns:</p> <ul> <li><em>dataSetName – </em>a name of a data set on which evaluation was carried out; there are three possible values: (i) <em>firstPages</em> refers to the first data set (<em>L<sub>D</sub></em>) that contains textual description of a company; (ii) <em>firstPageLabels</em> refers to the second data set (<em>L<sub>L</sub></em>) that involves link labels that were extracted from an index page; (iii) <em>aggregateDocument</em> refers to the third data set (<em>L<sub>B</sub></em>) that consists of a so-called big document;</li> <li><em>fmeasure</em> - the number of features that were taken into account during evaluation;</li> <li><em>method</em> - the name of function in the caret package;</li> <li><em>parameters</em> - the values of parameters received from a tuning phase of a given classification method;</li> <li><em>precision </em>– the value of method’s precision;</li> <li><em>recall </em>– the value of method’s recall;</li> <li><em>fmeasure</em> - the value of method’s F-measure; </li> <li><em>error</em> - the value of method’s error;</li> <li><em>acc </em>– the value of method’s.</li> </ul> <p><strong>Time processing statistics</strong></p> <p>All time processing statistics, like the performance statistics, are stored in cvs files. Each file corresponds to a particular machine learning method such as a file, "methodName-time.csv". All files cover the following columns:</p> <ul> <li><em>dataSetName – </em>a name of a data set on which evaluation was carried out; there are three possible values: (i) <em>firstPages</em> refers to the first data set (<em>L<sub>D</sub></em>) that contains textual description of a company; (ii) <em>firstPageLabels</em> refers to the second data set (<em>L<sub>L</sub></em>) that involves link labels that were extracted from an index page; (iii) <em>aggregateDocument</em> refers to the third data set (<em>L<sub>B</sub></em>) that consists of a so-called big document;</li> <li><em>featureNo</em> - the number of features that were taken into account during evaluation;</li> <li><em>method</em> - the name of function in the caret package;</li> <li><em>user</em> - user time elapsed for executing a <em>method</em> as an R process;</li> <li><em>system</em> - system time elapsed for executing a <em>method</em> as an R process;</li> <li><em>elapsed</em> - total time elapsed for executing a <em>method</em> as an R process.</li> </ul> <p>For more information about user, system and total elapsed time, please see documentation [3].</p> <p><strong>References</strong></p> <p>[1] https://cran.r-project.org/web/packages/caret/</p> <p>[2] https://topepo.github.io/caret/model-training-and-tuning.html</p> <p>[3] https://stat.ethz.ch/R-manual/R-devel/library/base/html/proc.time.htm</p>
Measuring Web Latency and Rendering Performance: Method, Tools & Longitudinal Dataset
<p>The dataset used in the paper entitled "Measuring Web Latency and Rendering Performance: Method, Tools & Longitudinal Dataset" published in IEEE Transactions for Network and Service Management. </p>
On the Use of Artificially Degraded Manuscripts for Quality Assessment of Readability Enhancement Methods - Dataset & Code
<p>This object contains the dataset and python code used for the paper:</p> <p>S. Brenner and R. Sablatnig. On the Use of Artificially Degraded Manuscripts for Quality Assessment of Readability Enhancement Methods. Accepted for OAGM Workshop 2019<strong>, </strong>Steyr, Austria.</p> <p>The dataset is a modified subset of the UCL Multispectral Processed Images of Parchment Damage Dataset (<a href="http://dx.doi.org/10.14324/000.ds.1469099">10.14324/000.ds.1469099</a>). The accompanying code documents how the modified version was created and how the evaluations described in the paper were performed.</p>
Data used in paper "A comparative study of calibration methods for low-cost ozone sensors in IoT platforms"
<p>Data used in paper "A comparative study of calibration methods for low-cost ozone sensors in IoT platforms", submitted for publication. The data consists of: (i) raw data from three nodes with four MICS 2614 metal-oxide ozone sensors deployed in Spain, summer 2017, and (ii) raw data of five alphasense OX-B431 and NO2-B43F electro-chemical sensors, four deployed in Italy and one in Austria, summers 2017 and 2018. Moreover, we have added the calibrated data using four machine learning methods: Multiple Linear Regression (MLR), K-Nearest Neighbors (KNN), Random Forest (RF) and Support Vector Regression (SVR).</p>
Dataset for "Reflectance spectra of seven lunar swirls examined by statistical methods: A space weathering study"
<p>This archive corresponds to the source code, raw data, and results described in the article "Reflectance spectra of seven lunar swirls examined by statistical methods: A space weathering study" by Chrbolková et al. (2019) published in Icarus journal. See AA_README.txt for more information.</p>
Dataset for "Method to retrieve cloud condensation nuclei number concentrations using lidar measurements"
<p>This repository contains the source data for the manuscript "<strong>Method to retrieve cloud condensation nuclei number concentrations using lidar measurements</strong>" published in <em>Atmospheric Measurement Techniques</em>. In situ measured data from five filed campaigns and corresponding theoretical simulated CCN number concentrations, lidar extinction and backscatter are included.</p>
Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method
<p>This dataset contains a GitHub repository containing all the data, analysis, Nextflow workflows and Jupyter notebooks to replicate the manuscript titled "Fast and accurate large multiple sequence alignments with a root-to-leaf regressive method".</p> <p>It also contains the Multiple Sequence Alignments (MSAs) generated and well as the main figures and tables from the manuscript.</p> <p>The repository is also available at GitHub (https://github.com/cbcrg/dpa-analysis) release `v1.2`.</p> <p>For details on how to use the regressive alignment algorithm, see the T-Coffee software suite (https://github.com/cbcrg/tcoffee).</p>
ArcGIS Map Packages and GIS Data for: A Geospatial Method for Estimating Soil Moisture Variability in Prehistoric Agricultural Landscapes, Gillreath-Brown et al. (2019)
<p><strong>ArcGIS Map Packages and GIS Data for Gillreath-Brown, Nagaoka, and Wolverton (2019)</strong></p> <p>**When using the GIS data included in these map packages, please cite all of the following:</p> <blockquote> <p>Gillreath-Brown, Andrew, Lisa Nagaoka, and Steve Wolverton. A Geospatial Method for Estimating Soil Moisture Variability in Prehistoric Agricultural Landscapes, 2019. PLoSONE 14(8):e0220457. <a href="http://doi.org/10.1371/journal.pone.0220457">http://doi.org/10.1371/journal.pone.0220457</a></p> <p>Gillreath-Brown, Andrew, Lisa Nagaoka, and Steve Wolverton. ArcGIS Map Packages for: A Geospatial Method for Estimating Soil Moisture Variability in Prehistoric Agricultural Landscapes, Gillreath-Brown et al., 2019. Version 1. Zenodo. <a href="https://doi.org/10.5281/zenodo.2572018">https://doi.org/10.5281/zenodo.2572018</a></p> </blockquote> <p><strong>OVERVIEW OF CONTENTS</strong></p> <p>This repository contains map packages for Gillreath-Brown, Nagaoka, and Wolverton (2019), as well as the raw digital elevation model (DEM) and soils data, of which the analyses was based on. The map packages contain all GIS data associated with the analyses described and presented in the publication. The map packages were created in ArcGIS 10.2.2; however, the packages will work in recent versions of ArcGIS. (Note: I was able to open the packages in ArcGIS 10.6.1, when tested on February 17, 2019). The primary files contained in this repository are:</p> <ul> <li>Raw DEM and Soils data <ul> <li>Digital Elevation Model Data (Map services and data available from U.S. Geological Survey, National Geospatial Program, and can be downloaded from the <a href="https://viewer.nationalmap.gov/basic/">National Elevation Dataset</a>) <ul> <li><strong>DEM_Individual_Tiles</strong>: Individual DEM tiles prior to being merged (1/3 arc second) from USGS National Elevation Dataset.</li> <li><strong>DEMs_Merged</strong>: DEMs were combined into one layer. Individual watersheds (i.e., Goodman, Coffey, and Crow Canyon) were clipped from this combined DEM. </li> </ul> </li> <li> Soils Data (Map services and data available from <a href="https://data.nal.usda.gov/dataset/natural-resources-conservation-service-web-soil-survey">Natural Resources Conservation Service Web Soil Survey</a>, U.S. Department of Agriculture) <ul> <li><strong>Animas-Dolores_Area_Soils</strong>: Small portion of the soil mapunits cover the northeastern corner of the Coffey Watershed (CW).</li> <li><strong>Cortez_Area_Soils</strong>: Soils for Montezuma County, encompasses all of Goodman (GW) and Crow Canyon (CCW) watersheds, and a large portion of the Coffey watershed (CW).</li> </ul> </li> </ul> </li> <li>ArcGIS Map Packages <ul> <li><strong>Goodman_Watershed_Full_SMPM_Analysis</strong>: Map Package contains the necessary files to rerun the SMPM analysis on the full Goodman Watershed (GW).</li> <li><strong>Goodman_Watershed_Mesa-Only_SMPM_Analysis</strong>: Map Package contains the necessary files to rerun the SMPM analysis on the mesa-only Goodman Watershed.</li> <li><strong>Crow_Canyon_Watershed_SMPM_Analysis</strong>: Map Package contains the necessary files to rerun the SMPM analysis on the Crow Canyon Watershed (CCW).</li> <li><strong>Coffey_Watershed_SMPM_Analysis</strong>: Map Package contains the necessary files to rerun the SMPM analysis on the Coffey Watershed (CW).</li> </ul> </li> </ul> <p>For additional information on contents of the map packages, please see see "Map Packages Descriptions" or open a map package in ArcGIS and go to "properties" or "map document properties."</p> <p><strong>LICENSES</strong></p> <p>Code: <a href="http://opensource.org/licenses/MIT">MIT</a> year: 2019 <br> Copyright holders: Andrew Gillreath-Brown, Lisa Nagaoka, and Steve Wolverton</p> <p><strong>CONTACT</strong></p> <p><strong>Andrew Gillreath-Brown, PhD Candidate, RPA</strong><br> <a href="https://anthro.wsu.edu/">Department of Anthropology</a>, Washington State University<br> <a href="mailto:andrew.brown1234@gmail.com">andrew.brown1234@gmail.com</a> – Email<br> <a href="https://andrewgillreathbrown.wordpress.com/">andrewgillreathbrown.wordpress.com</a> – Web</p>
Images of apples for the use of the Viola-Jones method. Data set no. 2 - grey scale.
<p>The database contains pictures of apples made at different angles, from different sides and containing different varieties. In this way, two bases of apple images were created (each database contains 1,100 images). This set is data set no. 2 - grey scale: processed images in shades of gray. The photos were prepared for the best possible detection process in the Viola-Jones method. These photo bases with apples can be used to teach machines to recognize specific varieties and count apples.</p>
A Reproducible Comparison of RSSI Fingerprinting Localization Methods Using LoRaWAN (datasets)
<p>The train/validation/test sets used in the study "<strong>A Reproducible Comparison of RSSI Fingerprinting Localization Methods Using LoRaWAN</strong>".</p> <p>Preprint: <a href="https://arxiv.org/abs/1908.05085">https://arxiv.org/abs/1908.05085</a></p> <p>Published paper: <a href="https://ieeexplore.ieee.org/document/8970177">https://ieeexplore.ieee.org/document/8970177</a></p> <p> </p> <p>The dataset used to create these sets was published in:</p> <p><a href="http://www.mdpi.com/2306-5729/3/2/13">http://www.mdpi.com/2306-5729/3/2/13</a></p> <p>The full dataset is available here:</p> <pre><a href="https://doi.org/10.5281/zenodo.1212478">https://doi.org/10.5281/zenodo.1212478</a> </pre> <p>The credit for the creation of the dataset goes to Aernouts, Michiel; Berkvens, Rafael; Van Vlaenderen, Koen and Weyn, Maarten.</p>
Supplementary material: Efficient in vivo screening method for the identification of C4 photosynthesis inhibitors based on cell suspensions of the single-cell C4 plant Bienertia sinuspersici
<p>Data described in Minges et al. (2019) Efficient <em>in vivo</em> screening method for the identification of C<sub>4</sub> photosynthesis inhibitors based on cell suspensions of the single-cell C<sub>4</sub> plant <em>Bienertia sinuspersici</em>. doi: <a href="https://doi.org/10.3389/fpls.2019.01350">10.3389/fpls.2019.01350</a></p> <p> </p>
Global trends and collaborations in electrochemical methods: a dataset on etching and deposition research
<p><span>This dataset supports the study "Electrochemical Etching vs. Electrochemical Deposition: A Comparative Bibliometric Analysis," which examines scientific publications on electrochemical etching and electrochemical deposition from 1970 to 2023. The dataset is derived from the Science Citation Index Expanded (SCIE) database and includes bibliometric information on publication trends, leading contributors, research areas, and keyword co-occurrences in both fields.</span></p>
Experiments regarding the inverse method extended with the integer hull ("RITPS")
<p>This archive contains the benchmarks regarding the evaluation of the inverse method extended with integer hull ("RITPS"):</p> <ul> <li>the version of IMITATOR: <span><a href="https://github.com/imitator-model-checker/imitator/releases/tag/v3.4.0-alpha2">v3.4.0-alpha2</a></span> (Cheese Durian)</li> <li>the source models</li> <li>the script to run all versions of CSMA/CD</li> <li>the expected results (in subdirectory "res")</li> </ul>
Datasets used in the benchmarking study of MR methods
<p>We conducted a benchmarking analysis of 16 summary-level data-based MR methods for causal inference with five real-world genetic datasets, focusing on three key aspects: type I error control, the accuracy of causal effect estimates, replicability, and power.</p> <p>The datasets used in the MR benchmarking study can be downloaded here:</p> <ol> <li>"dataset-GWASATLAS-negativecontrol.zip": the GWASATLAS dataset for evaluation of type I error control in confounding scenario (a): Population stratification</li> <li>"dataset-NealeLab-negativecontrol.zip": the Neale Lab dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-PanUKBB-negativecontrol.zip": the Pan UKBB dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-Pleiotropy-negativecontrol": the dataset used for evaluation of type I error control in confounding scenario (b): Pleiotropy;</li> <li>"dataset-familylevelconf-negativecontrol.zip": the dataset used for evaluation of type I error control in confounding scenario (c): Family-level confounders;</li> <li>"dataset_ukb-ukb.zip": the dataset used for evaluation of the accuracy of causal effect estimates;</li> <li>"dataset-LDL-CAD_clumped.zip": the dataset used for evaluation of replicability and power;</li> </ol> <p>Each of the datasets contains the following files:</p> <ol> <li> "Tested Trait pairs": the exposure-outcome trait pairs to be analyzed;</li> <li>"MRdat" refers to the summary statistics after performing IV selection (p-value < 5e-05) and PLINK LD clumping with a clumping window size of 1000kb and an r^2 threshold of 0.001.</li> <li>"bg_paras" are the estimated background parameters "Omega" and "C" which will be used for MR estimation in MR-APSS.</li> </ol> <p>Note:</p> <ol> <li>The formatted dataset after quality control can be accessible at our GitHub website (https://github.com/YangLabHKUST/MRbenchmarking).</li> <li>The details on quality control of GWAS summary statistics, formatting GWASs, and LD clumping for IV selection can be found on the MR-APSS software tutorial on the MR-APSS website (https://github.com/YangLabHKUST/MR-APSS).</li> <li>R code for running MR methods is also available at https://github.com/YangLabHKUST/MRbenchmarking.</li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.