Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo32/100

Digital rock dataset for use in machine learning research of permeability prediction

<p>Datasets and codes for the paper &quot;Hierarchical homogenization with deep-learning-based surrogate model for rapid estimation of effective permeability from digital rocks&quot;</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Exploring the configuration space of elemental carbon with empirical and machine learned interatomic potentials

<p>This dataset contains a vertical slice of the data used to generate the results found in the&nbsp;publication &quot;Exploring the configuration space of elemental carbon with empirical and machine learned interatomic potentials&quot;<br> It contains nested sampling input files and trajectory files for each potential studied, as well as the xml files and training data for the new potential, GAP-20U+gr.</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Experimental Data for: Machine learning enabled image analysis of time-temperature sensing colloidal arrays

<p>This dataset contains images of colloidal arrays functioning as time-temperature integrating, autonomous sensors. Each image shows a sensor consisting of multiple colloidal crystals with varying compositions of particles with different glass transition temperatures. Details on the composition and manufacturing procedures are explained in the corresponding publication.&nbsp;The data is organized into folders with different temperature setpoints. For each temperature, we investigated 10 samples. The name of the samples corresponds to their creation date. The name of the image files corresponds to their acquisition time. The first image in each sample folder was taken at the very start of the heating period. Hence, the heating time can be calculated by subtracting the start time from the acquisition time.</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

An Empirical Investigation into Learning Bug-Fixing Patches in the Wild via Neural Machine Translation

<p>Paper:&nbsp;An Empirical Investigation into&nbsp;Learning Bug-Fixing Patches in the Wild via Neural Machine Translation</p> <p>Authors:&nbsp;Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk</p> <p>Journal:&nbsp;TOSEM 2019 - ACM Transactions on Software Engineering and Methodology&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

WeatherPon: A Weather and Machine Learning-based Coupon Recommendation Mechanism in Digital Marketing

<p>These&nbsp;are the datasets that we used.</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Data for "Machine-learning-aided atomic structure identification of interfacial ionic hydrates from AFM images"

<p>Dataset for Neural Network training and testing of paper&nbsp;entitled &quot;Machine-learning-aided atomic structure identification of interfacial ionic hydrates from AFM images&quot; (<a href="https://doi.org/10.1093/nsr/nwac282">https://doi.org/10.1093/nsr/nwac282</a>).</p> <p>Each file contains&nbsp;named-dependent simulated AFM Images at&nbsp;different tip height and corresponding atomic structure file in POSCAR format. (See detailed description in the manuscript <a href="https://doi.org/10.1093/nsr/nwac282">https://doi.org/10.1093/nsr/nwac282</a>)</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Summary ouput data - Wasteaware Cities Benchmark Indicators - WABI 2023 - Global data analytics - Machine learning vs. Non-linear Regression

<p>This is the output&nbsp;dataset for the research publication &quot;<em>Socio-economic development drives solid waste management performance in cities: A global analysis using machine learning</em>&quot;. It features&nbsp;</p> <ul> <li>Metadata info used by R codes</li> <li>Summary of results for two modelling approaches (machine learning:&nbsp;Conditional random-forest and non-linear regression)</li> </ul> <p>The independent variables dataset&nbsp;analysed here refer to specific indicators of the WABI methodology (<a href="https://www.sciencedirect.com/science/article/pii/S0956053X14004905">https://www.sciencedirect.com/science/article/pii/S0956053X14004905</a>) that generates solid waste management and resource recovery profiles for cities. It was&nbsp;applied here for 40 cities around the world. The data input are available here: 10.5281/zenodo.7570174</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Relevant Datasets and Software Used for Paper "KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description"

<p>This repository contains relevant datasets and software&nbsp;used in a paper&nbsp;&quot;KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description&quot;. They are used to run the code of <em>KGML-xDTD&nbsp;</em>stored on <a href="https://github.com/chunyuma/KGML-xDTD">Github</a>&nbsp;and support the results of this paper.</p> <p><strong>About the datasets</strong></p> <p>1. <em>bkg_rtxkg2c_v2.7.3.tar.gz</em></p> <p>This tar.gz file contains three sub-folders: tsv_files, scripts, and relevant_dbs. The &quot;tsv_files&quot; sub-folder has the input files that the neo4j software uses. The &quot;scripts&quot; sub-folder contains a shell script with a relevant python script to construct&nbsp;the&nbsp;biomedical knowledge graph. The &quot;relevant_dbs&quot; sub-folder stores two auxiliary databases that <em>KGML-xDTD</em> needs to use.&nbsp;</p> <p>2. <em>indication_paths.yaml</em></p> <p>This file contains the <a href="https://sulab.github.io/DrugMechDB">DrugMechDB</a>&nbsp;MOA paths that we used to evaluate the predicted MOA paths by <em>KGML-xDTD.&nbsp;</em>It is downloaded from the official <a href="https://github.com/SuLab/DrugMechDB">GitHub repository</a> of DrugMechDB.</p> <p>3.&nbsp;<em>training_data.tar.gz</em></p> <p>This tar.gz file contains the processed training data of four data sources (e.g., <a href="https://mychem.info">MyChem</a>, <a href="https://lhncbc.nlm.nih.gov/ii/tools/SemRep_SemMedDB_SKR/SemMedDB_download.html">SemMedDB</a>, <a href="https://bioportal.bioontology.org/ontologies/NDFRT">NDF-RT</a>, <a href="https://unmtid-shinyapps.net/shiny/repodb/">RepoDB</a>) mentioned in the paper. These processed drug-disease pairs have been matched to the identifiers of biological entities used in our biomedical knowledge graph and respectively split into true positive (tp) sets and true negative (tn) sets. We also provide the names of these drug identifiers and disease identifiers under a sub-folder &quot;translated _to_name&quot;.</p> <p><strong>About the software</strong></p> <p><em>neo4j-community-3.5.26.tar.gz</em></p> <p>This tar.gz is the Neo4j community version 3.5.26 downloaded from <a href="https://neo4j.com/download-center/#community">Neo4j Download Center</a>. Although the&nbsp;newer versions are&nbsp;available, due to their big&nbsp;changes in the Neo4j setting that are not compatible with our scripts on Github, we provide the version that we used in our research. If you would like to use the newer version, modifications to our script will be required to import the biomedical knowledge graph into your local Neo4j database with the new setting.</p>

opencc-zeroJan 2023View details →
zenodo32/100

Tricycle Accident Prevention and Control System using Machine Learning Techniques

<p>Tricyle accident prevention and control system using a feed forward neural network.</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

CalcAMP: A new machine learning model for the accurate pre-diction of antimicrobial activity of peptides

<p>Datasets used for the publication:&nbsp;</p> <p>CalcAMP: A new machine learning model for the accurate prediction of antimicrobial activity of peptides</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Replication Data for: ``Toward machine learning-augmented, bathymetry-aware parameterizations of mesoscale eddy buoyancy fluxes across upwelling slope fronts''

<p>This dataset contains the Python&nbsp;scripts for&nbsp;training the Artificial Neural Networks (ANNs), the trained ANNs, configuration files&nbsp;for&nbsp;the reference 2D MITgcm&nbsp;simulations, and model&nbsp;outputs used in the paper.</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Screening of key risk SNPs for glioma based on machine learning algorithms

<p>Glioma is a common primary malignant brain tumor and is the most aggressive and lethal solid tumor, accounting for approximately 80% of all intracranial malignancies. Our aim was to screen key SNP&nbsp;by LASSO regression and random forest (a machine learning algorithm) and construct a model based on these SNP&nbsp;to predict the risk of glioma in Chinese Han population.</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Quantitative Assessment of the Impact of Future Land Use Changes on Flood Risk Using Remote Sensing, Machine Learning, and a Hydraulic Model

<p>&nbsp;</p> <p>The RF Machine learning code&nbsp;</p> <p>Topological, geomorphology, geology, metrological information of the Tajan watershed.</p> <p>Land use land cover images of the Tajan watershed</p> <p>River, transportation roads, villages map&nbsp;</p> <p>Global damage function datasets.</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

Assessing the antimicrobial Capacity of metal oxide nanomaterials using ZOI and MIC measurements by employing  Machine learning tools.

<p>Assessing the antimicrobial Capacity of metal and metal oxide&nbsp;nanomaterials (IONPs, AgNPs, ZnONPs) using ZOI and MIC measurements by employing Machine learning tools.</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

pywaterinfo dataset for master's dissertation: Updating a conceptual rainfall-runoff model based on radar observation and machine learning

<p>This forcings dataset is the output of the pywaterinfo (https://fluves.github.io/pywaterinfo/) read in of forcing data (rain and potential evapotranspiration).</p> <p>Code related to this dataset can be found here:&nbsp;https://github.com/olivierbonte/master_thesis</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

OpenEO dataset for master's dissertation: Updating a conceptual rainfall-runoff model based on radar observation and machine learning

<p>This dataset is the output of the <a href="https://openeo.org/">OpenEO</a>&nbsp;processing of satellite data&nbsp;(SAR backscatter and LAI).&nbsp;</p> <p>Code related to this dataset can be found <a href="https://github.com/olivierbonte/master_thesis">here</a></p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Streamflow Predictions using Machine Learning with Data Reformation

<p>Streamflow Predictions using Machine Learning with Data Reformation</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

A Machine Learning Approach to Visual Perception of Forest Trails for Mobile Robots

<p>This&nbsp;dataset is a part of the supplementary materials to the 2017 RAL&nbsp;<a href="https://ieeexplore.ieee.org/document/7358076">article</a>&nbsp;with the same title.</p> <blockquote> <p>A Machine Learning Approach to Visual Perception of Forest Trails for Mobile Robots</p> <p>IEEE Robotics and Automation Letters</p> <p>Alessandro Giusti, Jerome Guzzi, Dan Ciresan, Fang Lin He, Juan Pablo Rodriguez, Flavio Fontana, Matthias Faessler, Christian Forster, Jurgen Schmidhuber, Gianni A. Di Caro, Davide Scaramuzza, Luca Gambardella</p> </blockquote> <p>You can find more information on&nbsp;the&nbsp;<a href="http://bit.ly/perceivingtrails">project web page</a>&nbsp;(alessandrog@idsia.ch).</p> <p><strong>Dataset</strong></p> <p>Folders 001..010 contain the dataset used to train the networks. Folder 000 contains preliminary test data. Folders 011..014 contain data for testing the system.</p> <ul> <li>000 and 003 were shot with an handheld cellphone.</li> <li>001 and 002 were shot with 3 GOPRO Hero 3 cameras, fixed on the head with straps.</li> <li>004..014 were shot with 3 Bluefox cameras, fixed on a rigid helm (the same model and with the same lens as the camera mounted on the quadcopter).</li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo32/100

A Machine Learning based approach to osteoporosis classification: correlational and comparative analysis between Osseus and DXA exams

<p>The osseus dataset is composed of data from 505&nbsp;individuals who underwent the osseus triage and DXA exam at&nbsp;the University Hospital Onofre Lopes of Federal University of Rio Grande do Norte, Brazil. This dataset provides elementary data to analyze the prediction of changes in bone mineral density by Osseus&nbsp;using supervised classification algorithms. Supplementary file&nbsp;presents the dictionary used during the data analysis.</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

The machine learning based statistical emulators of GGCMI phase 2

<p>A statistical emulator with machine learning algorithm to reproduce the response of year-to-year variation of four crop yield to CO<sub>2</sub> (C), temperature (T), water (W) and nitrogen (N) perturbations defined in the Global Gridded Crop Model Intercomparison Project (GGCMI) phase 2 experiment.</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record