Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo28/100

A Comparison of Machine-Learning Assisted Optical and Thermal Camera Systems for Beehive Activity Counting

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

Dataset: Machine Learning Based Prediction of Polaron-Vacancy Patterns on the TiO$_2$(110) Surface

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

Machine learning potential for serpentines

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Global distribution and trends of Thyroid cancer and its machine learning prediction: Comprehensive findings and questions from global burden of disease 1990–2021

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Driving factors of HONO and their synergy in Hohhot, a semi-arid city in northern China: Insights from interpretable machine learning approaches

<p>Research data includes air pollutants observation data and meteorological parameters data.</p>

opencc-by-4.0Nov 2024View details →
zenodo28/100

Supplementary Data for "Predicting Thermodynamic Stability of Inorganic Compounds Using Ensemble Machine Learning Based on Electron Configuration"

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Dataset for Application of machine learning to choke performance

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Dataset [ref. paper "Predictive modeling of drivers' brake reaction time through machine learning methods"]

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2024View details →
zenodo28/100

Machine Learning Mid-Infrared Spectral Models for Predicting Modal Mineralogy of CI/CM Chondritic Asteroids and Bennu

<p>This is supporting data for the&nbsp;paper titled &quot;Machine Learning Mid-Infrared Spectral Models for Predicting Modal Mineralogy of CI/CM Chondritic Asteroids and Bennu&quot;.&nbsp;Figure S1 compares model performance of nonnegative LSMA and PLS from Pan et al. (2015).&nbsp;Tables S1 through S5 provide XRD data, metadata, MIR spectra using the laboratory conversion, MIR spectra using the OTES conversion and quantitative XRD results for Murchison meteorite.&nbsp;</p>

openother-openNov 2021View details →
zenodo28/100

Figure 1 from: Wilf P, Wing SL, Meyer HW, Rose JA, Saha R, Serre T, Cúneo NR, Donovan MP, Erwin DM, Gandolfo MA, González-Akre E, Herrera F, Hu S, Iglesias A, Johnson KR, Karim TS, Zou X (2021) An image dataset of cleared, x-rayed, and fossil leaves vetted to plant family for human and machine learning. PhytoKeys 187: 93-128. https://doi.org/10.3897/phytokeys.187.72350

Figure 1 Selected image pairs of confamilial extant and fossil (see Appendix 1) leaves from the dataset ABatesia floribunda Spruce ex. Benth. (Fabaceae), NCLC-W 6417, showing typical layout of a cleared-leaf slide with original annotations (other examples are cropped in this figure); source voucher Froes 12074, DS 291771 (at CAS), Amazonas, Brazil BFabaceae sp. CJ1, SGC-ICP-10173; Cerrejón mine, middle-late Paleocene of Guajira Peninsula, Colombia CCrataegus viridis L. (Rosaceae), NCLC-W 11951b; H. Meyer s/n (collected 1974, no other voucher), cultivated, California, USA DCrataegus copeana (Rosaceae), UCMP 3610; Florissant, late Eocene of Colorado, USA; H. Meyer photograph number 0420 ETetracentron sinense Oliv. (Trochodendraceae), S. Wing negative 71-002; E.H. Wilson 659, US 599036, Szechuan, China FZiziphoides flabellum (Trochodendraceae), USNM 560134; Mexican Hat, early Paleocene of Montana, USA GQuercus prinus L. (Fagaceae), NCLC-W 6137; H. Foster 8223, US 1730249, Florida, USA HFagopsis longifolia (Fagaceae), FLFO 003432A; Florissant, late Eocene of Colorado, USA IEucalyptus astringens (Maiden) Maiden (Myrtaceae), NCLC-W 10489; J.H. Maiden (9 November 1909), Western Australia, UC 437518 JEucalyptus frenguelliana (Myrtaceae), MPEF-Pb 2344; Laguna del Hunco, early Eocene of Chubut, Argentina KCercidiphyllum obtritum (Cercidiphyllaceae), DMNH 25061; Republic, early Eocene of Washington, USA LCercidiphyllum japonicum Siebold &amp; Zucc. ex J.J.Hoffm. &amp; J.H.Schult.bis (Cercidiphyllaceae), Axelrod cleared leaf 166; UCMP (no other voucher) MPlatanus racemosa Nutt. (Platanaceae), NCLC-H 6631; Handel s/n (collected 1985, no other voucher), California, USA NErlingdorfia montana (compound-leaved Platanaceae), DMNH 7642; Hell Creek Formation, Late Cretaceous of North Dakota, USA. Scale bars: centimeters as labeled (A, B, L, M); 1 cm when not labeled (C–K, N).

opencc-by-4.0Dec 2021View details →
zenodo28/100

OpenABC-D: A Large-Scale Dataset For Machine Learning Guided Integrated Circuit Synthesis

<p>Logic synthesis is a challenging and widely-researched combinatorial optimization problem during integrated circuit (IC) design. It transforms a high-level description of hardware in a programming language like Verilog into an optimized digital circuit netlist, a network of interconnected Boolean logic gates, that implements the function. Spurred by the success of ML in solving combinatorial and graph problems in other domains, there is growing interest in the design of ML-guided logic synthesis tools. Yet, there are no standard datasets or prototypical learning tasks defined for this problem domain. Here, we describe OpenABC-D,a large-scale, labeled dataset produced by synthesizing open source designs with a leading open-source logic synthesis tool and illustrate its use in developing, evaluating and benchmarking ML-guided logic synthesis. OpenABC-D has intermediate and final outputs in the form of 870,000 And-Inverter-Graphs (AIGs) produced from 1500 synthesis runs plus labels such as the optimized node counts, and de-lay. We define a generic learning problem on this dataset and benchmark existing solutions for it. The codes related to dataset creation and benchmark models are available athttps://github.com/NYU-MLDA/OpenABC.git.</p>

opencc-by-4.0Mar 2022View details →
zenodo28/100

data for "Validation of Rapid and Low-Cost Approach for the Delineation of Zone Management Based on Machine Learning Algorithms"

<p>These data were used in article &quot;Validation of Rapid and Low-Cost Approach for the Delineation of Zone Management Based on Machine Learning Algorithms&quot; (https://doi.org/10.3390/agronomy12010183)</p>

opencc-by-4.0Apr 2022View details →
zenodo28/100

Data and Scripts used in "Analysis of Relations Between Solar Activity, Cosmic Rays and Earth Climate Using Machine Learning Techniques"

<p>This archive contains the data in relation to the work:</p> <p>Analysis of relations between solar activity, cosmic rays and earth climate using machine learning techniques<br> B. Belen, U. M. Leloglu, and M. B. Demirkoz&nbsp;</p> <p>See README file for more details.</p>

opencc-by-4.0Apr 2022View details →
zenodo28/100

Improving the local climate zone classification with building height, imperviousness, and machine learning

<p>Dataset and python script for the Local Climate Zone classification analysis&nbsp;used in the study.&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo28/100

Training data set for inverse pyrolysis machine learning

<p>For dataset description see https://doi.org/10.5281/zenodo.6606247</p>

opencc-by-4.0Mar 2022View details →
zenodo28/100

Reproducibility data for tables in "A machine learning approach to portfolio pricing and risk management for high dimensional problems"

<p>This dataset contains all the information necessary for reproducing the tables in the paper&nbsp;&quot;A machine learning approach to portfolio pricing and risk management for high dimensional problems&quot;.</p> <p>The raw benchmark data can be found in the Zenodo dataset &quot;Benchmark and training data for replicating financial and insurance examples&quot; (https://zenodo.org/record/3837381). To the extend necessary, only summary data from that dataset is used in this dataset.</p> <p>The dataset includes a jupyter notebook file that explains what the different files contain, and provides sample code to analyze the information and reproduce the tables.</p>

opencc-by-4.0Jun 2022View details →
zenodo28/100

Accuracy of EEG Biomarkers in the Detection of Clinical Outcome in Disorders of Consciousness after Severe Acquired Brain Injury: Preliminary Results of a Pilot Study Using a Machine Learning Approach

<p>Dataset for the accepted publication: &quot;Accuracy of EEG Biomarkers in the Detection of Clinical Outcome in Disorders of Consciousness after Severe Acquired Brain Injury: Preliminary Results of a Pilot Study Using a Machine Learning Approach&quot;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo28/100

Data and code for "Machine learning can guide food security efforts when primary data is not available"

<p>Data and code repository for the paper &quot;Machine learning can guide food security efforts when primary data is not available&quot; by Giulia Martini, Alberto Bracci, Lorenzo Riches, Sejal Jaiswal, &nbsp;Matteo Corea, Jonathan Rivers, Arif Husain, and Elisa Omodei.</p>

opencc-by-4.0Aug 2022View details →
zenodo28/100

Replication Package of the paper "Machine Learning-based Test Selection for Simulation-based Testing of Self-driving Cars Software"

<p># Replication Package&nbsp;of the paper&nbsp;&quot;Machine Learning-based Test Selection for Simulation-based Testing of Self-driving Cars Software&quot;<br> ## SDC-Scissor (Self-Driving Car coSt-effeCtIve teSt SelectOR)</p> <p><br> ### Structure<br> - datasets: the datasets we used in our study for training and evaluating ML Models<br> - RQs: the extended results and analysis for each of the RQs mentioned in the paper &nbsp;</p> <p>More information on each topic can be found in the README file in the subfolders.<br> &nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo28/100

Prediction of creep failure time using machine learning

<p>Dataset from an elastoplastic element creep model with disorder from this publication:</p> <p>https://www.nature.com/articles/s41598-020-72969-6#author-information</p> <p>The data contains different disorder parameters, applied stresses and system sizes. Each sample is described by a csv file with four columns: The first one is a running index from zero to the number of rows in the file minus one, the second one is the number of relaxation steps performed (check the publication what this means), the third one is the index of the element being relaxed/failing and the fourth is a time increment. To get the total time, simply sum up the time increments. For information, please contact Soumyayajyoti Biswas.</p>

opencc-by-4.0Oct 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record