Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Similarity based Machine Learning
<p>Data and static code to produce SML and RML learning curves based on the paper "Improved decision making with similarity based machine learning"</p>
Dataset and scripts for publication "Property design of extruded magnesium-gadolinium alloys through machine learning"
<p>Data and scripts accompanying publication "Property design of extruded magnesium-gadolinium alloys through machine learning"</p>
Towards a practical framework to "ethics by design" data sharing and machine learning applications
<p>Responsible AI and data-driven applications can only be developed when teams integrate the ethical principles directly into the development process. An important prerequisite is the involvement of a diverse group of stakeholders who build and are affected by AI and data systems. We present a practical framework that helps teams build trustworthy AI systems and data strategies by combining expertise and training from philosophy, law, machine learning and design.</p>
Machine Learning-Assisted Discovery of Hidden States in Expanded Free Energy Space
<p>Collective variables (CVs) are crucial parameters in enhanced sampling calculations and strongly impact the quality of the obtained free energy surface. However, many existing CVs are unique to and dependent on the system they are constructed with, making the developed CV non-transferable to other systems. Herein, we develop a non-instructor-led deep autoencoder neural network (DAENN) for discovering general-purpose CVs. The DAENN is used to train a model by learning molecular representations upon unbiased trajectories that contain only the reactant conformers. The prior knowledge of nonconstraint reactants coupled with the here-introduced topology variable and loss-like penalty function are only required to make the biasing method able to expand its configurational (phase) space to unexplored energy basins. Our developed autoencoder is efficient and relatively inexpensive to use in terms of <em>a priori</em> knowledge, enabling one to automatically search for hidden CVs of the reaction of interest.</p>
Dataset for "Calibration of CAMS PM2.5 data over Hungary: A machine learning approach"
<p>This study utilized air quality information from eleven particular checking locales arranged all through Hungary. The informational index contains in-situ estimations of particulate matter with a width of 2.5 micrometers or less (PM2.5). Moreover, the dataset incorporates matching PM2.5 gauges from the Copernicus Climate Checking Administration (CAMS) model, which makes it a valuable asset for correlation and adjustment. The dataset incorporates various seasons, considering an exhaustive assessment of PM tainting under different climatic circumstances.</p>
Embedded machine learning to promote detection of unsafe environments
<p>Datasets used to train and validate the devices and model described in the paper "<strong>Embedded machine learning of IoT streams to promote early detection of unsafe environments</strong>"</p>
NGS data from: Deploying synthetic coevolution and machine learning to engineer protein-protein interactions
<p>Fine-tuning of protein-protein interactions occurs naturally through coevolution, but this process is difficult to recapitulate in the laboratory. We describe a synthetic platform for protein-protein coevolution that can isolate matched pairs of interacting muteins from complex libraries. This large dataset of coevolved complexes<span class="Apple-converted-space"> </span>drove a systems-level analysis of molecular recognition between Z domain-affibody pairs spanning a wide range of structures, affinities, cross-reactivities, and orthogonalities, and captured a broad spectrum of coevolutionary networks. Furthermore, we harnessed pre-trained protein language models to expand, <em>in silico</em>, the amino acid diversity of our coevolution screen, predicting remodeled interfaces beyond the reach of the experimental library. The integration of these approaches provides a means of generating protein complexes with diverse molecular recognition properties as tools for biotechnology and synthetic biology.</p>
Ensemble of optimised machine learning algorithms for predicting surface soil moisture content at global scale (v1.0)
<p>This study investigates the estimation of daily SSM using eight optimised ML algorithms and ten ensemble models (constructed via model bootstrap aggregating techniques and five-fold cross-validation). The algorithmic implementations were trained and tested using the international soil moisture network (ISMN) data collected from 1722 stations distributed across the World. </p>
Machine Learning Integrated High Quantum Yield Blue Light Carbon Dots for Real-time and On-site Detection of Cr(VI) in Groundwater and Drinking Water
<p>RGB和Kmeans提取后含有Cr(VI)水样的图像数据</p>
Dataset for publication: "Using explainable machine learning to interpret the effects of policies on air pollution: COVID-19 lockdown in London"
<p>This online repository offers supplementary datasets supporting the findings in the research article titled "Using explainable machine learning to interpret the effects of policies on air pollution: COVID-19 lockdown in London," published in Environmental Science & Technology. The dataset contains air quality data from different monitoring sites in London between 2016 and 2020 and weather data for the same period. Additionally, the dataset incorporates 136 features relating to London's Middle Layer Super Output Areas (MSOAs) in the year 2019, which can be used to identify the key factors contributing to the heterogeneous changes in air quality levels at different spatial locations during the pandemic. Please refer to the metadata file for detailed data sources and descriptions. </p>
Supplementary material – Sedimentary organic matter accumulation provinces in the Santos Basin, SW Atlantic: insights from multiple bulk proxies and machine learning analysis
<p>Table S1 - Supplementary material: a complete dataset of the bulk and isotopic composition of organic matter, pigments, biopolymers, and related indices obtained in surface sediments from the Santos Basin. </p>
BubbleML: A Multi-Physics Dataset and Benchmarks for Machine Learning
<p>In the field of phase change phenomena, the lack of accessible and diverse datasets suitable for machine learning (ML) training poses a significant challenge. Existing experimental datasets are often restricted, with limited availability and sparse ground truth data, impeding our understanding of this complex multi-physics phenomena. To bridge this gap, we present the <a href="https://github.com/HPCForge/BubbleML">BubbleML</a> Dataset which leverages physics-driven simulations to provide accurate ground truth information for various boiling scenarios, encompassing nucleate pool boiling, flow boiling, and sub-cooled boiling. This extensive dataset covers a wide range of parameters, including varying gravity conditions, flow rates, sub-cooling levels, and wall superheat, comprising 79 simulations. BubbleML is validated against experimental observations and trends, establishing it as an invaluable resource for ML research. Furthermore, we showcase its potential to facilitate exploration of diverse downstream tasks by introducing two benchmarks: (a) optical flow analysis to capture bubble dynamics, and (b) operator networks for learning temperature dynamics. The BubbleML dataset and its benchmarks serve as a catalyst for advancements in ML-driven research on multi-physics phase change phenomena, enabling the development and comparison of state-of-the-art techniques and models.</p>
Marine heatwave prediction using machine learning - Moana Project
<p>This datasets includes results of the PCA analysis used in the repository https://github.com/metocean/marineheatwave_ml_moana</p> <p> </p>
Hybrid Machine Learning Model for Ultra-Short-Term Wind Power Forecasting with Multi-Model Training Approach
<p>This is the core data code of the "<strong>Hybrid Machine Learning Model for Ultra-Short-Term Wind Power Forecasting with Multi-Model Training Approach".</strong></p>
Linear machine learning based force matching for amorphous silica: How close are the classical two-body potentials to ab initio calculations?
<p>Please later see our manuscript (in submission) for details.</p>
GPRChinaSPEI1km: High spatial resolution and century-long SPEI datasets for China from 1901 to 2020 generated by machine learning
<p>The high spatial resolution and century-long Standardized Precipitation Evapotranspiration Index (SPEI) dataset with a spatial resolution of 0.0083 degrees (~1 km) was spatially downscaled from the global SPEI data with a 0.5 degrees spatial resolution (https://spei.csic.es/database.html) based on machine learning integrated with high spatial resolution climatic and topographic variables. The 1-km SPEI datasets are across the land areas of China from January 1901 to December 2020, including 1-month, 3-month, 6-month and 12-month SPEIs. The unit of the data is 0.01. The dataset was evaluated using the root zone soil moisture and the historical drought events, and the evaluation indicated that the high spatial resolution SPEI dataset is reliable.</p> <p>Data Information: </p> <p>GPRChinaSPEI1km: High spatial resolution and century-long SPEI datasets over China from 1901 to 2020 generated by machine learning</p> <p>Publication: </p> <p><span>He, Q., Wang, M., Liu, K., & Wang, B. (2025). High-resolution Standardized Precipitation Evapotranspiration Index (SPEI) reveals trends in drought and vegetation water availability in China. <em>Geography and Sustainability</em>, <em>6</em>(2), 100228. https://doi.org/10.1016/j.geosus.2024.08.007</span></p> <p></p> <p>----------------------------------------------------data description---------------------------------------------</p> <p>This is a gridded dataset for the Standardized Precipitation Evapotranspiration Index (SPEI) at a spatial resolution of 1 km over the main terrestrial lands of China for each month during 1901-2020, which is generated using the Gaussian process regression (GPR) based on the Global SPEI database (https://spei.csic.es/database.html) integrated with high spatial resolution climatic and topographic variables. Four timescales of SPEI were generated: 1-month (SPEI-1), 3-month (SPEI-3), 6-month (SPEI-6) and 12-month (SPEI-12). The details are as follows:</p> <p>Region: China</p> <p>Temporal Extent: January 1901 to December 2020</p> <p>Spatial resolution: 0.0083° (~1 km)</p> <p>Temporal resolution: month</p> <p>Timescales: 1-month, 3-month, 6-month and 12-month</p> <p>Data format: GeoTIFF</p> <p>Unit: unitless (0.01)</p> <p>Geographic coordinate system: WGS 1984</p> <p>---------------------------------------------------dataset filename---------------------------------------------</p> <p>The file name specifically shows the data information.</p> <p>For example,</p> <p>“SPEI_1_2020_1.tif” means “1-month SPEI of January 2020”.</p> <p>“SPEI_3_2020_1.tif” means “3-month SPEI of January 2020”.</p> <p>All the file names are formatted in “SPEI_timescale_year_month”</p> <p>timescale: 1, 3, 6 and 12 indicate 1-month, 3-month, 6-month and 12-month, respectively</p> <p>year: from 1901 to 2020</p> <p>month: from 1 to 12</p> <p>--------------------------------------------------storage information-------------------------------------------</p> <p>The high-resolution SPEI dataset is stored in TIFF format using WGS 1984 coordinate system. The data type is int16 with a scale factor of 0.01. The nodata value is -32768. The dataset requires multiplication by 0.01 during application to obtain the actual value ranges.</p> <p>The data were compressed into .rar format every 10 years for each timescale SPEI.</p>
Capturing the interactions in the BaSnF$_4$ ionic conductor: a comparison of machine-learning potentials and polarizable force fields
<p>The current files contain all the original data for our work "Capturing the interactions in the BaSnF<sub>4 </sub>ionic conductor: a comparison of machine-learning potentials and polarizable force fields". </p>
Ensemble BLUP, Machine Learning, and Deep Learning Models Predict Maize Yield Better Than Each Model Alone.
<p>Data and scripts exploring ensembling strategies using the models developed in <a href="https://academic.oup.com/g3journal/advance-article/doi/10.1093/g3journal/jkad006/6982634">Kick et al., 2023</a> (see also <a href="https://zenodo.org/record/7401113">1</a>, <a href="https://zenodo.org/record/6916775">2</a>). Download all files to a single directory then run setup.sh or manually unzip using tar.</p> <p> </p> <table> <tbody> <tr> <td><strong>Filename</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>setup.sh</td> <td>Simple script that unzips zipped directories</td> </tr> <tr> <td>ext_data</td> <td>Reduced data from Kick et al. 2023</td> </tr> <tr> <td>ext_data_notebooks</td> <td>Contains python notebooks containing analysis and R markdown file containing visualization of results. Python and R data objects are written to allow results to be read in instead of re-generated.</td> </tr> <tr> <td>output</td> <td>Folder containing a placeholder file.</td> </tr> </tbody> </table> <p> </p> <p>This research used resources provided by the United States Department of Agriculture’s Agricultural Research Service (project number 5070-21000-041-000-D). The SCINet project of the USDA Agricultural Research Service (project number 0500-00093-001-00-D) was instrumental in the training of the models used in this work. In addition, we would like to acknowledge those presently and historically involved in generating data for the Genomes to Fields Initiative.</p> <p> </p> <p> </p> <p> </p>
Kinetics of N2 Release from Diazo Compounds: A Combined Machine Learning-Density Functional Theory Study
<p>Total potential (E) and Thermal correction to Gibbs Free Energies obtained using SMD/M06-2X/def2-TZVP//SMD/M06-2X/6-31G(d) level of theory in dichloroethane and Cartesian coordinates for all of the calculated structures.</p> <p>dataset</p> <p>Python Machine Learning Script</p> <p> </p>
Raw Data for: "Inorganic synthesis-structure maps in zeolites with machine learning and crystallographic distances"
<p>This repository contains all the raw data to reproduce the manuscript:</p> <p>D. Schwalbe-Koda et al. "Inorganic synthesis-structure maps in zeolites with machine learning and crystallographic distances". arXiv:2307.10935 (2023)</p> <p>The raw data should be used in combination with the code hosted on GitHub: <a href="https://github.com/dskoda/Zeolites-AMD">https://github.com/dskoda/Zeolites-AMD</a>.</p> <p><strong>Description of the data</strong></p> <p>The data in this link contains all necessary information to reproduce the manuscript. In combination with the code hosted on GitHub, it can be visualized and analyzed accordingly. The full description on the columns and results is available on the GitHub code.<br> The data files in this repository are:</p> <p>- `hparams_rnd_*.json`: results of the hyperparameter optimization of all classifiers studied in this work. The data was produced by randomly sampling the train-validation-test sets. In some cases, the data was normalized (`_norm_`), and the train set was kept `balanced` or `unbalanced`.<br> - `hyp_dm`: distance matrix of all hypothetical zeolites towards the known zeolites<br> - `hyp_predictions`: predictions of the synthesis conditions for all hypothetical zeolites<br> - `xgb_ensembles*`: pickle files containing the serialized ensemble models used in the evaluation of the data in this work. The models can be loaded with the `xgboost` Python package.</p> <p><strong>License</strong></p> <p>The data and all the content from this repository is distributed under the Creative Commons Attribution 4.0 (CC-BY 4.0)</p> <p>This work was produced under the auspices of the U.S. Department of Energy by Lawrence Livermore National Laboratory under Contract DE-AC52-07NA27344.</p> <p>Dataset released as: LLNL-MI-854709.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.