Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Clinical Dataset for the Paper 'Machine Learning Prediction of Treatment Response to Biological Disease-Modifying Antirheumatic Drugs in Rheumatoid Arthritis'
<p>This dataset accompanies the manuscript titled "Machine Learning Prediction of Treatment Response to Biological Disease-Modifying Antirheumatic Drugs in Rheumatoid Arthritis." It includes clinical data used for training and evaluating the machine learning models described in the paper. The dataset contains baseline clinical data of 154 RA patients who were treated with bDMARDs. The labels for remission, and effectiveness (remission and low disease activity) were applied after a 6-month follow-up based on EULAR criteria on DAS28ESR. The sustained effectiveness label indicates maintaining effectiveness within 6 months after initially achieving effectiveness.</p> <p><strong>Crossponder Authors:</strong></p> <ul> <li>Fatemeh Salehi (email: <a rel="noreferrer">fatemeh.salehihafshejni@fau.de</a>)</li> </ul>
VirMAD - Interpretable Machine Learning for Bacterial Defense and Viral Antidefense Pair Inference
Open the record for dataset details and reuse information.
Shear Sonic Prediction Using Supervised Machine Learning: Case Study Talang Akar Formation
<p>This material has presented on 2nd International Conference on Advanced Research in Engineering and Technology in October 25, 2023.</p>
Predicting critical transitions with surrogate data-based machine learning: Project
<p>This project accompanies the Github repository <a href="https://github.com/ZhiqinMa/surrogate_data_based_machine_learning">https://github.com/ZhiqinMa/surrogate_data_based_machine_learning</a>. It contains the model time series data that are used to train the machine learning algorithms, as well as all the process and result files. Please extract all <code>.rar</code> files to <code>G:/surrogate_data_based_machine_learning_v4.0.0</code> folder to run this project.</p>
Multi-Scale Computational Design of Metal-Organic Frameworks for Carbon Capture Using Machine Learning and Multi-Objective Optimization
<p>This repository contains CIF files for metal-organic frameworks and Grand canonical Monte Carlo (GCMC) simulation results for the article <em>Multi-Scale Computational Design of Metal-Organic Frameworks for Carbon Capture Using Machine Learning and Multi-Objective Optimization</em> by Zijun Deng and Lev Sarkisov.</p>
Machine learning on intraplate magmatism in NE China: Data and Code
<p>This version has a few errors. For the correct version, please see https://zenodo.org/records/12806006.</p>
Bridging the Language Gap: An Empirical Study of Bindings for Open Source Machine Learning Libraries in Software Package Ecosystems
<p>Replication package for the paper "Bridging the Language Gap: An Empirical Study of Bindings for Open Source Machine Learning Libraries in Software Package Ecosystems "</p>
Data for figures of manuscript entitled: "On the Sample Complexity of Quantum Boltzmann Machine Learning"
<p>The zip file contains the data for each of the plots in the figures in the manuscript: "On the Sample Complexity of Quantum Boltzmann Machine Learning." The preprint version of this article can be found on arXiv: https://arxiv.org/abs/2306.14969</p>
Data for Modeling Snow on Sea Ice using Physics Guided Machine Learning
Open the record for dataset details and reuse information.
Chest X-Ray Image Dataset: A Resource for Medical Diagnosis and Machine Learning
<p>The Chest X-Ray Image Dataset is an extensive collection designed to support medical research and the development of diagnostic tools for COVID-19 detection. It consists of two distinct classes: COVID-19 affected X-ray images and normal X-ray images of the chest area, each covering the full lungs. This dataset provides a diverse range of X-ray images, capturing the unique characteristics of both healthy and COVID-19 affected lungs, making it an invaluable resource for training and testing machine learning models in medical image classification and analysis.</p>
Satellite and Celestial Data for Machine Learning (SCD-ML)
<p>These datasets contains information about Kosmos 2514 satellite, and can be used for machine learning. The data comes from the <a href="https://www.juntadeandalucia.es/institutodeestadisticaycartografia/" target="_blank" rel="noopener">Institute of Statistics and Cartography of Andalusia (IECA)</a>. Then there is a detailed explanation:</p> <ul> <li><em><strong>satellite_data:</strong></em> This dataset includes ephemerides (precise satellite position and velocity data) about the satellite</li> <li><em><strong>sgdp4_celestial_data</strong></em>: This dataset contains ephemerides for the Kosmos 2514 satellite, positions of various celestial bodies in the solar system, and SGDP4 predictions for the satellite's position.</li> <li><em><strong>sequential_data_smj</strong></em>: This dataset includes sequences of 10 positions for the Sun, Moon, and Jupiter.</li> <li><em><strong>sequential_data_svmmj</strong></em>: Similar to the previous dataset, this one contains sequences of 10 positions for the Sun, Venus, Moon, Mars, and Jupiter. Each sequence also spans from the initial position to the SGDP4 predicted position of the satellite.</li> </ul> <p>The last two datasets, <em><strong>sequential_data_smj </strong></em>and <em><strong>sequential_data_svmmj</strong></em>, only provide sequences that are linked to the corresponding rows in <em><strong>sgdp4_celestial_data</strong></em>. They detail 10 positions of celestial bodies over the period between the initial position and the SGDP4 predicted position of the satellite.</p>
Linking satellites to genes with machine learning to estimate phytoplankton community structure from space
<p><strong>General description</strong></p> <p>The datasets presented in this repository have served in the development of a new ocean color algorithm to derive the relative cell abundance of seven phytoplankton groups (output of algorithm #1, called SOMRCA), as well as their contribution to total chlorophyll a (ChlaPG, output of algorithm #2, called SOMChlF) at the global scale using an omic-based marker: psbO. The outputs of the algorithm SOMChlF were compared to the HPLC-based definition of phytoplankton groups.</p> <p>All the details about this study are found in El Hourany, R., Pierella Karlusich, J., Zinger, L., Loisel, H., Levy, M., and Bowler, C.: Linking satellites to genes with machine learning to estimate phytoplankton community structure from space, Ocean Sci., 20, 217–239, https://doi.org/10.5194/os-20-217-2024, 2024.</p> <p>In "Tara_Oceans_psbO_dataset_Final.xlsx", it can be found the Tara Oceans' psbO metagenomic counts converted into relative cell abundance and Chlorophyll-a contribution for seven phytoplankton groups alongside satellite matchups. This dataset was used for algorithm development. In "Assets HPLC_SOMChlF.xlsx", the HPLC database was used as a comparison with Satellite-derived ChlaPG. Each data document presents a description sheet.</p> <p>In the following, the datasets used in this study are described.</p> <p><strong>Tara Oceans psbO metagenomic abundances</strong><br>The psbO gene is a single-copy gene in most eukaryotes and prokaryotes. We used psbO reads from the metagenomes generated by the Tara Oceans expedition as a proxy for phytoplankton relative cell abundance (see more details in Pierella Karlusich et al., 2023 Mol Ecol Res; https://doi.org/10.1111/1755-0998.13592).</p> <p>Among the 210 Tara Oceans stations, 145 stations sampled metagenomes in different ocean regimes from oligotrophic to eutrophic waters (Chl a from 0.01 to 10 mg m−3, median at 0.3 mg m−3) from 2009 to 2013. Seawater samples were filtered to differentiate five planktonic size fractions (0.22–3, 0.8–5, 5–20, 20–180, 180–2000 µm). <br>We retrieved the psbO read abundances from each Tara Oceans size-fractionated seawater sample from the Supplementary Material from Pierella Karlusich et al., 2023 Mol Ecol Res (https://www.ebi.ac.uk/biostudies/files/S-BSST761/psbO_mapping_against_Tara_Oceans_metagenomes.tsv).</p> <p>We used the psbO data to taxonomically differentiate seven phytoplankton groups: diatoms, dinoflagellates, green algae, haptophytes, pelagophytes, cryptophytes, and prokaryotes (cyanobacteria). The psbO read abundances of these seven groups are expressed as relative phytoplankton cell abundance (%) and their contribution to the Chlorophyll-a (Chla PG, in mg m-3). Phytoplankton that were not assigned to any of these seven groups (unclassified) represented less than 5 % of the total relative cell abundance among all size classes. To obtain a single value of relative cell abundance per station, we pooled the five size fractions into a single aggregated sample. For Chl a content estimation, we used a conversion via size-dependent weights (see formula 1 in El Hourany et al., 2024).</p> <p>There are two levels of information derived from the molecular dataset: relative abundance of psbO reads as a proxy for relative cell abundance and the fraction of Chl a that each group represents. Both types of information have different implications. Chl a is often used as a proxy for biomass, which is a relevant parameter for energy and matter fluxes (e.g., food webs, biogeochemical cycles). At the same time, cell abundance corresponds to species abundance for unicellular organisms, which is an important measure for inferring community assembly processes.<br> <br><strong>Satellite Matchups</strong><br>We used ocean color products from the GlobColour project (R2019, full archive reprocessed, 2020) to retrieve satellite matchups for the psbO-derived abundances. These products were constructed by merging data from various satellite sensors: Sea-viewing Wide Field-of-view Sensor (SeaWiFS), Moderate Resolution Imaging Spectroradiometer (MODIS), Visible Infrared Imaging Radiometer Suite (VIIRS), Medium Resolution Imaging Spectrometer (MERIS), and Ocean and Land Colour Instrument (OLCI).</p> <p>We used 16 GlobColour products as inputs to retrieve the phytoplankton community structure: chlorophyll a concentration (Chl a, product name: CHL1-AVW), remote sensing reflectances (Rrs) at 11 wavelengths (412, 443, 469, 490, 510, 531, 547, 555, 620, 645, and 670 nm), light attenuation coefficient at 490 nm (Kd490), photosynthetically available radiation (PAR), normalized fluorescence light height (NFLH), and particulate backscattering at 443 nm (bbp). These products have daily and 4 km spatiotemporal resolution. In addition, we used the Climate Change Initiative (CCI) sea surface temperature (SST) product at 4 km resolution and daily frequency distributed by the Copernicus Marine Services (CMEMS) portal.</p> <p><strong>HPLC datasets</strong><br>To compare satellite-derived phytoplankton group Chla fractions' distribution (outputs of the algorithm named SOMChlF) with more conventional DPA-based products, we compiled a global HPLC dataset regrouping 12 000 HPLC observations from several HPLC datasets between 1997 and 2014. This HPLC dataset was collocated with the SOMChlF-based ChlaPG. This dataset depicts the abundance of the pigments most widely used to identify major phytoplankton groups: fucoxanthin (Fuco), peridinin (Perid), alloxanthin (Allo), zeaxanthin (Zea), chlorophyll b (Chl b), 19-hexanoyloxyfucoxanthin (19HF), and 19-butanoyloxyfucoxanthin (19BF).</p> <p>Diagnostic pigments were used to estimate the Chl a fraction for each phytoplankton group, namely diatoms, dinoflagellates, haptophytes, green algae, cryptophytes, pelgophytes, and prokaryotes. The Chl a fraction per group is expressed by</p> <p>HPLC-based ChlaPG = Chla in-situ · DP · α / Sum (DP · α) where "α" is a coefficient associated with a diagnostic pigment (DP) for a specific PG.</p> <p>All the details are found in El Hourany, R., Pierella Karlusich, J., Zinger, L., Loisel, H., Levy, M., and Bowler, C.: Linking satellites to genes with machine learning to estimate phytoplankton community structure from space, Ocean Sci., 20, 217–239, https://doi.org/10.5194/os-20-217-2024, 2024.</p> <p> </p>
Implementation of a Diabetes Status Prediction Application Using a Machine Learning Algorithm Approach
Open the record for dataset details and reuse information.
Gaussian approximation of dispersion potentials for efficient featurization and machine-learning predictions of metal–organic frameworks
<p>Scripts and data for the publication</p>
Data for the manuscript "Spatially resolved uncertainties for machine learning potentials"
<p>This repository accompanies the manuscript "Spatially resolved uncertainties for machine learning potentials" by E. Heid, J. Schörghuber, R. Wanzenböck, and G. K. H. Madsen. The following files are available:</p> <ul> <li> <p><code>mc_experiment.ipynb</code> is a Jupyter notebook for the Monte Carlo experiment described in the study (artificial model with only variance as error source).</p> </li> </ul> <ul> <li> <p><code>aggregate_cut_relax.py</code> contains code to cut and relax boxes for the water active learning cycle.</p> </li> <li> <p><code>data_t1x.tar.gz</code> contains reaction pathways for 10,073 reactions from a subset of the Transition1x dataset, split into training, validation and test sets. The training and validation sets contain the indices 1, 2, 9, and 10 from a 10-image nudged-elastic band search (40k datapoints), while the test set contains indices 3-8 (60k datapoints). The test set is ordered according to the reaction and index, i.e. rxn1_index3, rxn1_index4, [...] rxn1_index8, rxn2_index3, [...].</p> </li> <li> <p><code>data_sto.tar.gz</code> contains surface reconstructions of SrTiO3, randomly split into a training and validation set, as well as a test set.</p> </li> <li> <p><code>data_h2o.tar.gz</code> contains:</p> <ul> <li> <p><code>full_db.extxyz</code>: The full dataset of 1.5k structures.</p> </li> <li> <p><code>iter00_train.extxyz</code> and <code>iter00_validation.extxyz</code>: The initial training and validation set for the active learning cycle.</p> </li> <li> <p>the subfolders in the folders <code>random</code>, and <code>uncertain</code>, and <code>atomic</code> contain the training and validation sets for the random and uncertainty-based (local or atomic) active learning loops.</p> </li> </ul> </li> </ul>
Machine Learning Models for cfMeDIP data from Shen et al.
<p>Contains raw markdowns and knit markdown with plots for machine learning analyses in Shen et al, "<strong>Highly sensitive tumor detection and classification using methylome analysis of plasma cell-free DNA".</strong></p>
Tools for prediction of tumor heterogeneity by a machine learning approach
<p>The package includes R codes and datasets. We applied three classification algorithms: Support Vector Machine, Random Forest, and Naïve Bayes. Datasets include tab-delimited files of mutation and gene expression profile for stomach cancer. </p>
A 1-km resolution monthly mean air temperature (Ta) dataset across the Tibetan Plateau during 2001-2015, related with the article "Mapping monthly air temperature in the Tibetan Plateau from MODIS data based on machine learning methods".
<p>We present a 1-km resolution monthly mean air temperature (Ta) dataset across the Tibetan Plateau from 2001 to 2015. It ranges from 25°-45°N, 70°-105°E, covering a total area of ~7,045,000 km2. To develop this dataset, 10 machine learning algorithms were applied to 11 environmental variables derived from Moderate Resolution Imaging Spectroradiometer (MODIS) data, Shuttle Radar Topography Mission (SRTM) digital elevation model (DEM) data and topographic index data. The best model generated by Cubist algorithm was finally selected to calculate monthly mean Ta, and achieved an overall accuracy of RMSE= 1.00 °C and MAE= 0.73 °C. To get details of this dataset, please refer to the manuscript "Mapping monthly air temperature in the Tibetan Plateau from MODIS data based on machine learning methods". This Ta dataset provides spatially continuous coverage compared with station observed data, and has much higher accuracy and spatial resolution than reanalysis datasets, making it a useful dataset for climate change and environmental studies in the Tibetan Plateau.</p> <p>Xu Y., Knudby A., Shen Y., Liu Y., Mapping monthly air temperature in the Tibetan Plateau from MODIS data based on machine learning methods. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2018, 11(2): 345-354. (DOI: 10.1109/jstars.2017.2787191).</p> <p>The Ta dataset is provided in ENVI standard format. The coordinate system is WGS84 Geographic Coordinate System.</p>
yaleemmlc/admissionprediction: Predicting hospital admission at emergency department triage using machine learning - Data and Scripts
<p>First release for PLOS One. Please cite original paper for all research using this dataset.</p>
Intermediate data objects from running the machine learning code for Shen et al, Nature, 2018
<p>These are the RData objects of the processed cfMeDIP data that were used to run the machine learning analyses in "Sensitive tumour detection and classification using plasma cell-free DNA methylomes", Nature, 2018. This archive also includes the models we generated, and training and validation data matrices that people can use to fit new models and evaluate performance. These objects are to be used with the scripts and markdown at doi: 10.5281/zenodo.1242697 . The ReadMe contains descriptions. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.