Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Digital rock dataset for use in machine learning research of permeability prediction
<p>Datasets and codes for the paper "Hierarchical homogenization with deep-learning-based surrogate model for rapid estimation of effective permeability from digital rocks"</p>
Exploring the configuration space of elemental carbon with empirical and machine learned interatomic potentials
<p>This dataset contains a vertical slice of the data used to generate the results found in the publication "Exploring the configuration space of elemental carbon with empirical and machine learned interatomic potentials"<br> It contains nested sampling input files and trajectory files for each potential studied, as well as the xml files and training data for the new potential, GAP-20U+gr.</p>
Experimental Data for: Machine learning enabled image analysis of time-temperature sensing colloidal arrays
<p>This dataset contains images of colloidal arrays functioning as time-temperature integrating, autonomous sensors. Each image shows a sensor consisting of multiple colloidal crystals with varying compositions of particles with different glass transition temperatures. Details on the composition and manufacturing procedures are explained in the corresponding publication. The data is organized into folders with different temperature setpoints. For each temperature, we investigated 10 samples. The name of the samples corresponds to their creation date. The name of the image files corresponds to their acquisition time. The first image in each sample folder was taken at the very start of the heating period. Hence, the heating time can be calculated by subtracting the start time from the acquisition time.</p>
An Empirical Investigation into Learning Bug-Fixing Patches in the Wild via Neural Machine Translation
<p>Paper: An Empirical Investigation into Learning Bug-Fixing Patches in the Wild via Neural Machine Translation</p> <p>Authors: Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk</p> <p>Journal: TOSEM 2019 - ACM Transactions on Software Engineering and Methodology </p>
WeatherPon: A Weather and Machine Learning-based Coupon Recommendation Mechanism in Digital Marketing
<p>These are the datasets that we used.</p>
Data for "Machine-learning-aided atomic structure identification of interfacial ionic hydrates from AFM images"
<p>Dataset for Neural Network training and testing of paper entitled "Machine-learning-aided atomic structure identification of interfacial ionic hydrates from AFM images" (<a href="https://doi.org/10.1093/nsr/nwac282">https://doi.org/10.1093/nsr/nwac282</a>).</p> <p>Each file contains named-dependent simulated AFM Images at different tip height and corresponding atomic structure file in POSCAR format. (See detailed description in the manuscript <a href="https://doi.org/10.1093/nsr/nwac282">https://doi.org/10.1093/nsr/nwac282</a>)</p> <p> </p>
Summary ouput data - Wasteaware Cities Benchmark Indicators - WABI 2023 - Global data analytics - Machine learning vs. Non-linear Regression
<p>This is the output dataset for the research publication "<em>Socio-economic development drives solid waste management performance in cities: A global analysis using machine learning</em>". It features </p> <ul> <li>Metadata info used by R codes</li> <li>Summary of results for two modelling approaches (machine learning: Conditional random-forest and non-linear regression)</li> </ul> <p>The independent variables dataset analysed here refer to specific indicators of the WABI methodology (<a href="https://www.sciencedirect.com/science/article/pii/S0956053X14004905">https://www.sciencedirect.com/science/article/pii/S0956053X14004905</a>) that generates solid waste management and resource recovery profiles for cities. It was applied here for 40 cities around the world. The data input are available here: 10.5281/zenodo.7570174</p>
Relevant Datasets and Software Used for Paper "KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description"
<p>This repository contains relevant datasets and software used in a paper "KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description". They are used to run the code of <em>KGML-xDTD </em>stored on <a href="https://github.com/chunyuma/KGML-xDTD">Github</a> and support the results of this paper.</p> <p><strong>About the datasets</strong></p> <p>1. <em>bkg_rtxkg2c_v2.7.3.tar.gz</em></p> <p>This tar.gz file contains three sub-folders: tsv_files, scripts, and relevant_dbs. The "tsv_files" sub-folder has the input files that the neo4j software uses. The "scripts" sub-folder contains a shell script with a relevant python script to construct the biomedical knowledge graph. The "relevant_dbs" sub-folder stores two auxiliary databases that <em>KGML-xDTD</em> needs to use. </p> <p>2. <em>indication_paths.yaml</em></p> <p>This file contains the <a href="https://sulab.github.io/DrugMechDB">DrugMechDB</a> MOA paths that we used to evaluate the predicted MOA paths by <em>KGML-xDTD. </em>It is downloaded from the official <a href="https://github.com/SuLab/DrugMechDB">GitHub repository</a> of DrugMechDB.</p> <p>3. <em>training_data.tar.gz</em></p> <p>This tar.gz file contains the processed training data of four data sources (e.g., <a href="https://mychem.info">MyChem</a>, <a href="https://lhncbc.nlm.nih.gov/ii/tools/SemRep_SemMedDB_SKR/SemMedDB_download.html">SemMedDB</a>, <a href="https://bioportal.bioontology.org/ontologies/NDFRT">NDF-RT</a>, <a href="https://unmtid-shinyapps.net/shiny/repodb/">RepoDB</a>) mentioned in the paper. These processed drug-disease pairs have been matched to the identifiers of biological entities used in our biomedical knowledge graph and respectively split into true positive (tp) sets and true negative (tn) sets. We also provide the names of these drug identifiers and disease identifiers under a sub-folder "translated _to_name".</p> <p><strong>About the software</strong></p> <p><em>neo4j-community-3.5.26.tar.gz</em></p> <p>This tar.gz is the Neo4j community version 3.5.26 downloaded from <a href="https://neo4j.com/download-center/#community">Neo4j Download Center</a>. Although the newer versions are available, due to their big changes in the Neo4j setting that are not compatible with our scripts on Github, we provide the version that we used in our research. If you would like to use the newer version, modifications to our script will be required to import the biomedical knowledge graph into your local Neo4j database with the new setting.</p>
Tricycle Accident Prevention and Control System using Machine Learning Techniques
<p>Tricyle accident prevention and control system using a feed forward neural network.</p>
CalcAMP: A new machine learning model for the accurate pre-diction of antimicrobial activity of peptides
<p>Datasets used for the publication: </p> <p>CalcAMP: A new machine learning model for the accurate prediction of antimicrobial activity of peptides</p>
Replication Data for: ``Toward machine learning-augmented, bathymetry-aware parameterizations of mesoscale eddy buoyancy fluxes across upwelling slope fronts''
<p>This dataset contains the Python scripts for training the Artificial Neural Networks (ANNs), the trained ANNs, configuration files for the reference 2D MITgcm simulations, and model outputs used in the paper.</p>
Screening of key risk SNPs for glioma based on machine learning algorithms
<p>Glioma is a common primary malignant brain tumor and is the most aggressive and lethal solid tumor, accounting for approximately 80% of all intracranial malignancies. Our aim was to screen key SNP by LASSO regression and random forest (a machine learning algorithm) and construct a model based on these SNP to predict the risk of glioma in Chinese Han population.</p>
Quantitative Assessment of the Impact of Future Land Use Changes on Flood Risk Using Remote Sensing, Machine Learning, and a Hydraulic Model
<p> </p> <p>The RF Machine learning code </p> <p>Topological, geomorphology, geology, metrological information of the Tajan watershed.</p> <p>Land use land cover images of the Tajan watershed</p> <p>River, transportation roads, villages map </p> <p>Global damage function datasets.</p>
Assessing the antimicrobial Capacity of metal oxide nanomaterials using ZOI and MIC measurements by employing Machine learning tools.
<p>Assessing the antimicrobial Capacity of metal and metal oxide nanomaterials (IONPs, AgNPs, ZnONPs) using ZOI and MIC measurements by employing Machine learning tools.</p>
pywaterinfo dataset for master's dissertation: Updating a conceptual rainfall-runoff model based on radar observation and machine learning
<p>This forcings dataset is the output of the pywaterinfo (https://fluves.github.io/pywaterinfo/) read in of forcing data (rain and potential evapotranspiration).</p> <p>Code related to this dataset can be found here: https://github.com/olivierbonte/master_thesis</p>
OpenEO dataset for master's dissertation: Updating a conceptual rainfall-runoff model based on radar observation and machine learning
<p>This dataset is the output of the <a href="https://openeo.org/">OpenEO</a> processing of satellite data (SAR backscatter and LAI). </p> <p>Code related to this dataset can be found <a href="https://github.com/olivierbonte/master_thesis">here</a></p>
Streamflow Predictions using Machine Learning with Data Reformation
<p>Streamflow Predictions using Machine Learning with Data Reformation</p>
A Machine Learning Approach to Visual Perception of Forest Trails for Mobile Robots
<p>This dataset is a part of the supplementary materials to the 2017 RAL <a href="https://ieeexplore.ieee.org/document/7358076">article</a> with the same title.</p> <blockquote> <p>A Machine Learning Approach to Visual Perception of Forest Trails for Mobile Robots</p> <p>IEEE Robotics and Automation Letters</p> <p>Alessandro Giusti, Jerome Guzzi, Dan Ciresan, Fang Lin He, Juan Pablo Rodriguez, Flavio Fontana, Matthias Faessler, Christian Forster, Jurgen Schmidhuber, Gianni A. Di Caro, Davide Scaramuzza, Luca Gambardella</p> </blockquote> <p>You can find more information on the <a href="http://bit.ly/perceivingtrails">project web page</a> (alessandrog@idsia.ch).</p> <p><strong>Dataset</strong></p> <p>Folders 001..010 contain the dataset used to train the networks. Folder 000 contains preliminary test data. Folders 011..014 contain data for testing the system.</p> <ul> <li>000 and 003 were shot with an handheld cellphone.</li> <li>001 and 002 were shot with 3 GOPRO Hero 3 cameras, fixed on the head with straps.</li> <li>004..014 were shot with 3 Bluefox cameras, fixed on a rigid helm (the same model and with the same lens as the camera mounted on the quadcopter).</li> </ul>
A Machine Learning based approach to osteoporosis classification: correlational and comparative analysis between Osseus and DXA exams
<p>The osseus dataset is composed of data from 505 individuals who underwent the osseus triage and DXA exam at the University Hospital Onofre Lopes of Federal University of Rio Grande do Norte, Brazil. This dataset provides elementary data to analyze the prediction of changes in bone mineral density by Osseus using supervised classification algorithms. Supplementary file presents the dictionary used during the data analysis.</p>
The machine learning based statistical emulators of GGCMI phase 2
<p>A statistical emulator with machine learning algorithm to reproduce the response of year-to-year variation of four crop yield to CO<sub>2</sub> (C), temperature (T), water (W) and nitrogen (N) perturbations defined in the Global Gridded Crop Model Intercomparison Project (GGCMI) phase 2 experiment.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.