Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
A Comparison of Machine-Learning Assisted Optical and Thermal Camera Systems for Beehive Activity Counting
Open the record for dataset details and reuse information.
Dataset: Machine Learning Based Prediction of Polaron-Vacancy Patterns on the TiO$_2$(110) Surface
Open the record for dataset details and reuse information.
Machine learning potential for serpentines
Open the record for dataset details and reuse information.
Global distribution and trends of Thyroid cancer and its machine learning prediction: Comprehensive findings and questions from global burden of disease 1990–2021
Open the record for dataset details and reuse information.
Driving factors of HONO and their synergy in Hohhot, a semi-arid city in northern China: Insights from interpretable machine learning approaches
<p>Research data includes air pollutants observation data and meteorological parameters data.</p>
Supplementary Data for "Predicting Thermodynamic Stability of Inorganic Compounds Using Ensemble Machine Learning Based on Electron Configuration"
Open the record for dataset details and reuse information.
Dataset for Application of machine learning to choke performance
Open the record for dataset details and reuse information.
Dataset [ref. paper "Predictive modeling of drivers' brake reaction time through machine learning methods"]
Open the record for dataset details and reuse information.
Machine Learning Mid-Infrared Spectral Models for Predicting Modal Mineralogy of CI/CM Chondritic Asteroids and Bennu
<p>This is supporting data for the paper titled "Machine Learning Mid-Infrared Spectral Models for Predicting Modal Mineralogy of CI/CM Chondritic Asteroids and Bennu". Figure S1 compares model performance of nonnegative LSMA and PLS from Pan et al. (2015). Tables S1 through S5 provide XRD data, metadata, MIR spectra using the laboratory conversion, MIR spectra using the OTES conversion and quantitative XRD results for Murchison meteorite. </p>
Figure 1 from: Wilf P, Wing SL, Meyer HW, Rose JA, Saha R, Serre T, Cúneo NR, Donovan MP, Erwin DM, Gandolfo MA, González-Akre E, Herrera F, Hu S, Iglesias A, Johnson KR, Karim TS, Zou X (2021) An image dataset of cleared, x-rayed, and fossil leaves vetted to plant family for human and machine learning. PhytoKeys 187: 93-128. https://doi.org/10.3897/phytokeys.187.72350
Figure 1 Selected image pairs of confamilial extant and fossil (see Appendix 1) leaves from the dataset ABatesia floribunda Spruce ex. Benth. (Fabaceae), NCLC-W 6417, showing typical layout of a cleared-leaf slide with original annotations (other examples are cropped in this figure); source voucher Froes 12074, DS 291771 (at CAS), Amazonas, Brazil BFabaceae sp. CJ1, SGC-ICP-10173; Cerrejón mine, middle-late Paleocene of Guajira Peninsula, Colombia CCrataegus viridis L. (Rosaceae), NCLC-W 11951b; H. Meyer s/n (collected 1974, no other voucher), cultivated, California, USA DCrataegus copeana (Rosaceae), UCMP 3610; Florissant, late Eocene of Colorado, USA; H. Meyer photograph number 0420 ETetracentron sinense Oliv. (Trochodendraceae), S. Wing negative 71-002; E.H. Wilson 659, US 599036, Szechuan, China FZiziphoides flabellum (Trochodendraceae), USNM 560134; Mexican Hat, early Paleocene of Montana, USA GQuercus prinus L. (Fagaceae), NCLC-W 6137; H. Foster 8223, US 1730249, Florida, USA HFagopsis longifolia (Fagaceae), FLFO 003432A; Florissant, late Eocene of Colorado, USA IEucalyptus astringens (Maiden) Maiden (Myrtaceae), NCLC-W 10489; J.H. Maiden (9 November 1909), Western Australia, UC 437518 JEucalyptus frenguelliana (Myrtaceae), MPEF-Pb 2344; Laguna del Hunco, early Eocene of Chubut, Argentina KCercidiphyllum obtritum (Cercidiphyllaceae), DMNH 25061; Republic, early Eocene of Washington, USA LCercidiphyllum japonicum Siebold & Zucc. ex J.J.Hoffm. & J.H.Schult.bis (Cercidiphyllaceae), Axelrod cleared leaf 166; UCMP (no other voucher) MPlatanus racemosa Nutt. (Platanaceae), NCLC-H 6631; Handel s/n (collected 1985, no other voucher), California, USA NErlingdorfia montana (compound-leaved Platanaceae), DMNH 7642; Hell Creek Formation, Late Cretaceous of North Dakota, USA. Scale bars: centimeters as labeled (A, B, L, M); 1 cm when not labeled (C–K, N).
OpenABC-D: A Large-Scale Dataset For Machine Learning Guided Integrated Circuit Synthesis
<p>Logic synthesis is a challenging and widely-researched combinatorial optimization problem during integrated circuit (IC) design. It transforms a high-level description of hardware in a programming language like Verilog into an optimized digital circuit netlist, a network of interconnected Boolean logic gates, that implements the function. Spurred by the success of ML in solving combinatorial and graph problems in other domains, there is growing interest in the design of ML-guided logic synthesis tools. Yet, there are no standard datasets or prototypical learning tasks defined for this problem domain. Here, we describe OpenABC-D,a large-scale, labeled dataset produced by synthesizing open source designs with a leading open-source logic synthesis tool and illustrate its use in developing, evaluating and benchmarking ML-guided logic synthesis. OpenABC-D has intermediate and final outputs in the form of 870,000 And-Inverter-Graphs (AIGs) produced from 1500 synthesis runs plus labels such as the optimized node counts, and de-lay. We define a generic learning problem on this dataset and benchmark existing solutions for it. The codes related to dataset creation and benchmark models are available athttps://github.com/NYU-MLDA/OpenABC.git.</p>
data for "Validation of Rapid and Low-Cost Approach for the Delineation of Zone Management Based on Machine Learning Algorithms"
<p>These data were used in article "Validation of Rapid and Low-Cost Approach for the Delineation of Zone Management Based on Machine Learning Algorithms" (https://doi.org/10.3390/agronomy12010183)</p>
Data and Scripts used in "Analysis of Relations Between Solar Activity, Cosmic Rays and Earth Climate Using Machine Learning Techniques"
<p>This archive contains the data in relation to the work:</p> <p>Analysis of relations between solar activity, cosmic rays and earth climate using machine learning techniques<br> B. Belen, U. M. Leloglu, and M. B. Demirkoz </p> <p>See README file for more details.</p>
Improving the local climate zone classification with building height, imperviousness, and machine learning
<p>Dataset and python script for the Local Climate Zone classification analysis used in the study. </p>
Training data set for inverse pyrolysis machine learning
<p>For dataset description see https://doi.org/10.5281/zenodo.6606247</p>
Reproducibility data for tables in "A machine learning approach to portfolio pricing and risk management for high dimensional problems"
<p>This dataset contains all the information necessary for reproducing the tables in the paper "A machine learning approach to portfolio pricing and risk management for high dimensional problems".</p> <p>The raw benchmark data can be found in the Zenodo dataset "Benchmark and training data for replicating financial and insurance examples" (https://zenodo.org/record/3837381). To the extend necessary, only summary data from that dataset is used in this dataset.</p> <p>The dataset includes a jupyter notebook file that explains what the different files contain, and provides sample code to analyze the information and reproduce the tables.</p>
Accuracy of EEG Biomarkers in the Detection of Clinical Outcome in Disorders of Consciousness after Severe Acquired Brain Injury: Preliminary Results of a Pilot Study Using a Machine Learning Approach
<p>Dataset for the accepted publication: "Accuracy of EEG Biomarkers in the Detection of Clinical Outcome in Disorders of Consciousness after Severe Acquired Brain Injury: Preliminary Results of a Pilot Study Using a Machine Learning Approach"</p> <p> </p>
Data and code for "Machine learning can guide food security efforts when primary data is not available"
<p>Data and code repository for the paper "Machine learning can guide food security efforts when primary data is not available" by Giulia Martini, Alberto Bracci, Lorenzo Riches, Sejal Jaiswal, Matteo Corea, Jonathan Rivers, Arif Husain, and Elisa Omodei.</p>
Replication Package of the paper "Machine Learning-based Test Selection for Simulation-based Testing of Self-driving Cars Software"
<p># Replication Package of the paper "Machine Learning-based Test Selection for Simulation-based Testing of Self-driving Cars Software"<br> ## SDC-Scissor (Self-Driving Car coSt-effeCtIve teSt SelectOR)</p> <p><br> ### Structure<br> - datasets: the datasets we used in our study for training and evaluating ML Models<br> - RQs: the extended results and analysis for each of the RQs mentioned in the paper </p> <p>More information on each topic can be found in the README file in the subfolders.<br> </p>
Prediction of creep failure time using machine learning
<p>Dataset from an elastoplastic element creep model with disorder from this publication:</p> <p>https://www.nature.com/articles/s41598-020-72969-6#author-information</p> <p>The data contains different disorder parameters, applied stresses and system sizes. Each sample is described by a csv file with four columns: The first one is a running index from zero to the number of rows in the file minus one, the second one is the number of relaxation steps performed (check the publication what this means), the third one is the index of the element being relaxed/failing and the fourth is a time increment. To get the total time, simply sum up the time increments. For information, please contact Soumyayajyoti Biswas.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.