Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo32/100

Distillation of crop models to learn plant physiology theories using machine learning

<p>Full Paper available for download at LINK</p>

opencc-by-4.0Mar 2019View details →
zenodo32/100

F I G U R E 3 in Machine learning for image based species identification

F I G U R E 3 Basic architecture of convolutional neural network (CNN). CNNs are comprised of one or more convolutional layers followed by one or more fully connected layers

opennotspecifiedDec 2018View details →
zenodo32/100

F I G U R E 4 Top-5 in Machine learning for image based species identification

F I G U R E 4 Top-5 Classification error rates of ImageNet Visual Recognition Challenge. In 2012, for the first time a deep neural network architecture (AlexNet) won the challenge

opennotspecifiedDec 2018View details →
zenodo32/100

F I G U R E 1 in Machine learning for image based species identification

F I G U R E 1 Typical human and computer vision pipeline for species identification. The machine learning platform takes in an image and outputs the confidence scores for a predefined set of classes

opennotspecifiedDec 2018View details →
zenodo32/100

Dataset for: Sky pixel detection in outdoor imagery using an adaptive algorithm and machine learning.

<p>The data presented in this article is related to the research article entitled ``Sky pixel detection in outdoor imagery using an adaptive algorithm and machine learning.&quot; \citep{Nice2019UC}.</p> <p>The dataset consists of a trained Inception V3 neural network model as well as the configuration files to train the neural network and run the inferences. The dataset also contains two sets of outdoor imagery (from Skyfinder and Google Street View) used to train the neural network and validate the sky pixel detection system in the linked article. The original images are included as well as rescaled imagery used to train the neural network, and sky masks used for validation.</p>

opencc-by-4.0Feb 2019View details →
zenodo32/100

Dataset for Machine Learning-Based Prediction and Optimization of As-Extruded Viability in Extrusion-Based 3D Bioprinting

<p>The dataset supports the findings presented in the paper "Machine learning-based prediction and optimization framework for as-extruded cell viability in extrusion-based 3D bioprinting."</p> <ul> <li>Sodium alginate viscosity data: "alg_i1g_viscosity_data.zip"</li> <li>Cross Power Law parameter fitting results: "alg_i1g_viscosity_fittings.zip"</li> <li>Rheological stability measurement: "alg_i1g_contact_angle_data.zip"</li> <li>OpenFOAM simulation results: "alg_i1g_simulation_data.zip"</li> <li>Post-extrusion cell viability results: "cell_viability_data.zip"</li> </ul> <p><strong>Keywords:</strong> 3D bioprinting; cell viability; shear stress; numerical analysis; machine learning; alginate<br><br></p> <p><strong>Code availability statement</strong><br>The scripts used for data analysis, machine learning models, and numerical simulations in this study are available on GitHub at:&nbsp;<a href="https://github.com/KORINZ/in-silico-bioink-viability-prediction">https://github.com/KORINZ/in-silico-bioink-viability-prediction</a></p>

restrictedcc-by-4.0Aug 2024View details →
zenodo32/100

Datasets for "Scalable interpolation of satellite altimetry data with probabilistic machine learning"

<p>Elevation (radar freeboard and sea-level anomaly) fields from CryoSat-2, Sentinel-3A, and Sentinel-3B, over the period December 1st 2018 - April 30th 2019. These data were processed for the Arctic domain using the European Space Agency's Grid Processing on Demand (GPOD) service. Processing follows the steps outlined in Lawrence et al., 2021 (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.asr.2019.10.011" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.asr.2019.10.011</a>). These data are provided at along-track, 5 km and 50 km resolution, where gridded data follow the EASE grid definition (<a href="https://doi.org/10.3390/ijgi1010032">https://doi.org/10.3390/ijgi1010032</a>).</p> <p>These data were used to develop the open-source Python programming library GPSat (https://github.com/CPOMUCL/GPSat), which uses local Gaussian Process models to perform scalable interpolation of non-stationary satellite altimetry data. The 'Source_data.xlsx' file contains the data corresponding to figures in the published Nature Communications article 'Scalable interpolation of satellite altimetry data with probabilistic machine learning'.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Data and machine-learning model for fcc-FeHx at the CMB conditions

<p>Data including melting temperatures and sound velocities. The machine-learning model was trained by DeePMD-kit. The initial configuration for two-phase coexistence simulations is provided.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Unsupervised Machine Learning and Cepstral Analysis with 4D-STEM for Characterizing Complex Microstructures of Metallic Alloys

<p>Raw 4D-STEM data of Ni50Ti26Hf20Al4 used for analysis in the publication "Unsupervised Machine Learning and Cepstral Analysis with 4D-STEM for Characterizing Complex Microstructures of Metallic Alloys". Datasets were collected using the electron microscope pixel array detector (EMPAD) with a Themis Z STEM. Custom python scripts used for data analysis are available upon request to one of the corresponding authors.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

VacSol-ML(ESKAPE) Machine Learning DataSet

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo32/100

Prevendo Aprovação de Empréstimos Utilizando Algoritmos de Machine Learning

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo32/100

MD simulations files for: Enhanced Sampling of Biomolecular Slow Conformational Transitions Using Adaptive Sampling and Machine Learning

<p>Here's a rephrased version of the README file:</p> <p>#### Enhanced Sampling of Biomolecular Slow Conformational Transitions Using Adaptive Sampling and Machine Learning</p> <p>**Authors:** Mingyuan Zhang, Hao Wu, Yong Wang</p> <p>This repository contains the official implementation for the paper "Enhanced Sampling of Biomolecular Slow Conformational Transitions Using Adaptive Sampling and Machine Learning" by Mingyuan Zhang, Hao Wu, and Yong Wang. Included are all trajectories from our MD simulations in the form of PLUMED COLVAR files, as well as all analysis scripts and files needed to replicate the results and figures presented in both the main text and Supporting Information (SI) of the paper.</p> <p>The paper features two examples: Ala2 and Ala10. For each, we have organized all the associated simulation files as they were during our automated simulation pipeline. The directory structure is the same for both examples. Here, we use Ala2, found in the `Ala2` folder, as an example:</p> <p>### Key Components</p> <p>- **Automated Pipeline Implementation:** The pipeline is implemented in `Ala2/7-adaptive-40ps/ala2.ipynb`. This implementation is ready to use once all required packages are installed, and gmx/gmx_mpi/plumed are callable within the notebook. After configuring the environment and specifying parameters like `gpu_id`, `ntomp`, and `n_sim` according to your hardware, running the blocks will replicate the entire pipeline.</p> <p>- **Analysis Scripts:** The scripts to replicate the results or figures from the main text or SI are organized in three files: `Ala2/7-adaptive-40ps/AdaptiveSamplingAnalysis.ipynb`, `Ala2/7-adaptive-40ps/compare_with_msm.ipynb`, and `Ala2/7-adaptive-40ps/opes/COLVAR/analysis.ipynb`.</p> <p>### Directory Structure</p> <p>Under the `Ala2` main directory, there are seven subdirectories:</p> <p>- **`Ala2/1-topol/`**: Contains files generated during system construction, including the final Gromacs topology file `topol.top`, which is necessary for running the automated simulation script.</p> <p>- **`Ala2/2-em/`, `Ala2/3-nvt/`, `Ala2/4-npt/`**: These directories store files generated during energy minimization and NVT/NPT equilibration. The `Ala2/4-npt/npt.gro` file is required to run the automated simulation script.</p> <p>- **`Ala2/mdp/`**: Contains all mdp files used, including `Ala2/mdp/md_detail.mdp`, which is necessary for running the automated simulation script.</p> <p>- **`Ala2/7-adaptive-40ps/`**: Contains all simulation and analysis scripts, along with files required to replicate the study related to the automated pipeline.</p> <p>&nbsp; 1. **`Ala2/7-adaptive-40ps/CV/`**: Stores all COLVAR files from adaptive sampling simulations.<br>&nbsp;&nbsp;<br>&nbsp; 2. **`Ala2/7-adaptive-40ps/opes/`**: Contains all files related to OPES simulations, including raw data for the final FES plots found in `Ala2/7-adaptive-40ps/opes/COLVAR/`. The script for replicating OPES and FES estimation figures is located in `Ala2/7-adaptive-40ps/opes/COLVAR/analysis.ipynb`.<br>&nbsp;&nbsp;<br>&nbsp; 3. **`Ala2/7-adaptive-40ps/figures/`**: Includes all original figures from the main text and SI, saved at 600 dpi.<br>&nbsp;&nbsp;<br>&nbsp; 4. **`Ala2/7-adaptive-40ps/traj_and_dat/`**: Stores all PLUMED `*.dat` files for the `DRIVER` utility in adaptive sampling simulations, a topology file `input.pdb` for PLUMED `MOLINFO`, and a topology file `seed_ref.pdb` for MDAnalysis adaptive sampling seed `*.gro` generation. Note that all `*.xtc` files from adaptive sampling were deleted to reduce the package size.<br>&nbsp;&nbsp;<br>&nbsp; 5. **Seed Index Files:** Seed indices for each round are stored as `Ala2/7-adaptive-40ps/round{i}_seed.txt`, necessary for figure replication.<br>&nbsp;&nbsp;<br>&nbsp; 6. **Automated Pipeline Notebook:** Implemented in `Ala2/ala2.ipynb`. Ensure that all imported packages are installed and gromacs (both gmx and gmx_mpi)/plumed can be called within the Jupyter notebook.<br>&nbsp;&nbsp;<br>&nbsp; 7. **Adaptive Sampling Analysis:** Scripts for analyzing adaptive sampling trajectories are found in `Ala2/AdaptiveSamplingAnalysis.ipynb`. This notebook contains scripts to replicate all figures related to adaptive sampling.<br>&nbsp;&nbsp;<br>&nbsp; 8. **MSM Comparison:** Analysis scripts for MSM comparison are located in `Ala2/compare_with_msm.ipynb`. This notebook contains scripts to replicate figures used for MSM/OPES comparison.</p> <p>- **`Ala2/8-adaptive-400ps/`**: Contains all simulation files (except xtc) for an additional adaptive sampling dataset computed for MSM comparison.</p> <p>### Contact Information</p> <p>We are continuing to test and improve the pipeline, so a tutorial is not yet available. Please feel free to reach out with any questions related to the implementation via email at mingyuanzhang@zju.edu.cn or by raising an issue on our GitHub page: https://github.com/yongwangCPH/papers/tree/main/2024/ALICE.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Dataset for "Machine Learning-Aided High Throughput Examination of Block Copolymer Processing Conditions"

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo32/100

Predicting Weather Disruptions for the ICC Champions Trophy 2025 in Pakistan Using Machine Learning and Data Analytics

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo32/100

First global machine learning model to predict solar flare impact on Earth's ionosphere

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo32/100

Estimation of landslide volume by machine learning and remote sensing techniques in Himalayan Regions

<p>This dataset was prepared for the development of a machine learning model so as to use this model to estimate the potential landslide volume in the Himalayan Regions, especially the Gyirong. This area serves as the only land route between China and Nepal, suffering from severe landslides each year and causing losses of human lives and properties. Therefore, developing an effective model for reliable estimation of landslide volume is of much significance for local people and decision-makers.</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Dataset for the research paper "Automated machine learning in research – a literature review"

<p>This repository contains the literature used in the research paper "Automated Machine Learning in Research &ndash; A Literature Review."</p> <p>The four BibTeX files contain the following collections of literature:</p> <table> <tbody> <tr> <td><strong>File</strong></td> <td><strong>Description</strong></td> <td><strong>Method</strong></td> </tr> <tr> <td><strong>01_Initial_papers.bib</strong></td> <td>385 Initial papers</td> <td>Keyword search in scientific databases Scopus and Web of Science</td> </tr> <tr> <td><strong>02_Primary_papers.bib</strong></td> <td>267 Primary papers</td> <td>Removing duplicate articles and filtering for full articles (i.e., conference and journal papers)</td> </tr> <tr> <td><strong>03_Possibly_relevant_papers.bib</strong></td> <td>54 Possibly relevant papers</td> <td>Identifying possibly relevant papers through abstract scan and application of inclusion criteria</td> </tr> <tr> <td><strong>04_Relevant_papers.bib</strong></td> <td>49 Relevant papers</td> <td>Inclusion of relevant papers after full-text analysis and backward and forward searches</td> </tr> </tbody> </table> <p>The following inclusion criteria were used to identify the 54 possibly relevant papers during the abstract scan:</p> <ul> <li>The abstract must mention autoML or related concepts such as low-code ML, meta-learning, or automated hyperparameter tuning and their potential use in research.</li> <li>The abstract must mention reproducibility or related concepts such as transparency or explainability in the context of autoML or the related concepts mentioned above</li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo32/100

A multi-species benchmark for training and validating large scale mass spectrometry proteomics machine learning models

<p>This is a de novo sequencing benchmark dataset derived from nine<br>publicly available mass spectrometry datasets. There are two versions<br>of the benchmark: main and balanced. The balanced version randomly<br>eliminates some spectra associated with some species in order to<br>create a smaller, more evenly balanced dataset. Also provided are two<br>zip files containing the raw data as well as intermediate results.<br>Details about how the benchmark was created are provided in an<br><a href="../records/13653420">associated zenodo release</a>, which contains the source code as well as a<br>manuscript describing the benchmark.</p> <p>This release fixes a bug that incorrectly detected shared peptides <br>between different species. It also includes the annotated spectra in <br>mzSpecLib format.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Code and Data for "Global Surface Eddy Mixing Ellipses: Spatio-temporal Variability and Machine Learning Prediction" By Jing et al. Submitted to Journal of Geophysical Research: Oceans.

<p>This repository contains the code and data for the study of "Global Surface Eddy Mixing Ellipses: Spatio-temporal Variability and Machine Learning Prediction" By Jing et al. Submitted to Journal of Geophysical Research: Oceans.</p> <p>Specifically, this repository contains the following items:&nbsp;</p> <p>(1) The codes needed for assessing the representation and&nbsp; prediction skills of Random Forest (RF) and Convolutional Neural Network (CNN) models.&nbsp;</p> <p>(2) Original and normalized data to run these codes.</p> <p>(3) &nbsp;Code here is built on early work from our laboratory (Guan et al., 2022; Zhang et al., 2023), though great modifications have been made tailored to our scientific question.</p> <div>[1] Guan, W., Chen, R., Zhang, H., Yang, Y., &amp; Wei, H. (2022). Seasonal surface eddy mixing in the Kuroshio Extension: Estimation and machine learning prediction. Journal of Geophysical Research: Oceans, 127 (3), e2021JC017967.</div> <div>[2]&nbsp;Zhang, G., Chen, R., Li, X., Li, L., Wei, H., &amp; Guan, W. (2023). Temporal variability of&nbsp;global surface eddy diffusivities: Estimates and machine learning prediction. Journal&nbsp;of Physical Oceanography, 53 (7), 1711&ndash;1730.</div>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Figure data for "Machine Learning For Early Dynamic Prediction Of Functional Outcome After Stroke"

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record