Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,075

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,075 results for “ML”

Learn how ShareScore rates datasets ↗
zenodo40/100

SPEA TESTBED TEST DATA FOR ML DEVELOPMENT

<p>The dataset files were generated at 3 different temperature with the improved ATE system developed in WP3 of the MET4FOF Project.&nbsp;</p> <p>The data are available for ML purposes and metrological investigation.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Prior C - posterior_statistics_using_ml

<p>Training set data (Prior C) as used in&nbsp;</p> <p>Hansen. T.M. and Finlay, C.C., Use of machine learning to estimate properties of the posterior distribution in probabilistic inverse problems - an application to airborne EM data. JGR - Solid Earth, &nbsp;20/10/2022.</p> <p>doi:10.1029/2022JB024703<br> &nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Prior A - posterior_statistics_using_ml

<p>Training set data (Prior A) as used in&nbsp;</p> <p>Hansen. T.M. and Finlay, C.C., Use of machine learning to estimate properties of the posterior distribution in probabilistic inverse problems - an application to airborne EM data. JGR - Solid Earth, &nbsp;20/10/2022.</p> <p>doi:10.1029/2022JB024703<br> &nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Open MatSci ML Toolkit - DGL Graphs for OpenCatalyst IS2RE Task (OC20)

<p>Dataset&nbsp;accompanying the release of the <a href="https://github.com/IntelLabs/matsciml">Open MatSciML Toolkit</a>, an open source software for development graph neural networks on the OpenCatalyst project using the Deep Graph Library (DGL).&nbsp;</p> <p>For more details about the Open MatSci ML Toolkit, check the associated <a href="https://github.com/IntelLabs/matsciml">open-source repository</a>&nbsp;and <a href="https://arxiv.org/abs/2210.17484">paper</a>.</p> <p>Compressed files ~8GB with uncompressed file being ~80&nbsp;GB.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

PYFOREST-ML-Predictions

<p>Results data from the Master&#39;s Capstone project PYFOREST. Predictions of deforestation under different policy scenarios are provided. See attached metadata for additional information, or refer to the&nbsp;<a href="https://github.com/cp-PYFOREST">GitHub project repository.</a>&nbsp;<br> &nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

TableS1 of ML-based predictive gut microbiome analysis for health assessment

<p>Complete list of species associated with the COVID and Control cohort (ANOVA F score &gt; 1). Comparison with respect to the original set of species used by Gupta <em>et al.</em>, and ANOVA-F values are reported.</p>

opencc-by-4.0Sep 2023View details →
ClinicalTrials.gov40/100

Study of Secukinumab With 2 mL Pre-filled Syringes

ClinicalTrials.gov study NCT02748863. IPD Sharing: UNDECIDED. Countries: 11. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo36/100

Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Fire-D: Analysis and ML-Ready NASA-Centric Remote Sensing of Wildfire and Smoke

<p>Earth science remote sensing imagery is rich in structural and spectral information, making such data an ideal platform for benchmarking for a broad range of machine learning (ML) tasks, from pattern retrieval to physics-informed classification to anomaly detection to transfer learning. Nevertheless, the utility of Earth science remote sensing data remains largely unexplored by the broader ML community. Our goal is to bridge this gap and bring a rich variety of multisource multi-resolution Earth image data to a wider range of ML researchers who are non-experts in remote sensing, thereby increasing the utility and societal impact of such data products. In particular, motivated by the emerging wildfire crisis, we present radiometrically and geometrically calibrated radiance data from airborne and orbital instruments from the National Aeronautics and Space Administration (NASA), the National Oceanic and Atmospheric Administration (NOAA), and the Korean Meteorological Administration (KMA).</p> <p>Given the scarce occurrence of wildfires and complex spatio-temporal dependencies in radiance data, these datasets are especially well suited for benchmarking unsupervised and self-supervised learning tasks both on images and non-Euclidean objects. Our experiments on these datasets indicate that contrastive learning and transfer learning algorithms can capture the structures of views and scenes, map pixel space of multi-sensor imagery to a high-level embedding space for further downstream tasks, and facilitate more cohesive integration of the state-of-the-art ML approaches into wildfire risk analytics.</p> <p>All NASA-based observations are freely usable under the <a href="https://science.data.nasa.gov/license/">Creative Commons Zero License</a>.There are also no restrictions on the use of <a href="https://registry.opendata.aws/noaa-goes/">GOES Data</a>.&nbsp;<a href="https://registry.opendata.aws/noaa-gk2a-pds/">GK2A data</a> are also open data without any restrictions on its use.<br><br>For the Planet data, we cannot not share the Radiances, but all masks within this dataset are freely usable with no restrictions.</p> <p>&nbsp;</p> <p>Use:</p> <p>On the data input, input geometrically and radiometrically calibrated radiance data has been pulled from various NASA, NOAA, Planet, and KMA archives. For instruments that have multiple different spatial resolutions within their spectral bands (GOES and GK2A), all bands have been resampled to the lowest collective spatial resolution.</p> <p>Geometric and radiometric calibration has been done by the science data processing pipelines of the various missions, and would not need to be done by anyone else looking to curate the same data. Further information for each instrument can be found in each of the publicly available Level-1 algorithm theoretical basis documents (ATBDs)</p> <p>All input and label data have been put in GeoTiff format. Each band is in a separate raster band and each scene is in a separate GeoTiff file. Label files and input files are in separate tar files, labeled respectively, and the file names match for input and labels, with the exception of an additional .fire and .smoke in the respective label filenames and subfolders.<br><br>The <a href="https://www.earthdata.nasa.gov/about/esdis/esco/standards-practices/geotiff">GeoTiff</a> data format natively contains geolocation metadata internally, and can be interfaced with via C/C++/Python <a href="https://gdal.org/en/stable">GDAL</a> packages, or other python packages that wrap GDAL, like <a href="https://rasterio.readthedocs.io/en/stable/">rasterio</a> and <a href="https://corteva.github.io/rioxarray/stable/">rioxarray</a> . The documentation for <a href="https://nicks-personal-organization-2.gitbook.io/sit-fuse">SIT-FUSE</a> , the package with which the labels were generated, also has examples on how to read and interface with various data formats, including GeoTiffs. Lastly, this data can be interfaced with using Geographic Information Systems (GIS), like the free and open-source <a href="https://qgis.org/">QGIS</a>.</p> <p>An example of programmatic data access and usage can be found in the dataset's associated <a href="https://github.com/Fire-D-Dataset/FIRE-D">GitHub repository</a>.&nbsp;</p> <p>A working example using data from this repository for ML tasks is available <a href="https://drive.google.com/drive/folders/16aJO6LhrxJ3gsWoTU9BNN3hsb8W0refG?usp=sharing">here</a>.</p> <p>Timing information can be found in the file names, which all use the standard formats from the various instruments' L1B datasets.</p> <p>V2 includes additional GOES-18 radiance data and associated smoke and fire labels for the recent LA fires (Palisades and Eaton fires in January of 2025).</p> <p>V3 provides a reorganization of all data, and an inclusion of improved and additional data from airborne and satellite platforms in 2019, associated with this study: https://arxiv.org/pdf/2501.15343 .&nbsp;</p> <p>V4 provides additional AVIRIS-C Radiances and fixes the spatial range of the GOES-17 radiances to match that of the associated labels. The AVIRIS-C radiances are split across 5 tar files, ordered temporally - all associated labels are in a single tar file.</p> <p><br>Current fire coverage includes:</p> <ul> <li>2019: Williams Flats, Sheridan, Horsefly, and Mosquito (US)</li> <li>2022: Uljin Forest Fire (S. Korea; largest fire on record in S. Korea)</li> <li>2025: Palisades and Eaton Fires (US)</li> </ul> <p>Additional data for the 2025 Palisades and Eaton fires from the TEMPO instrument is currently being validated and will be released in a V4 shortly.</p> <p>Croissant file for dataset metadata specification is also included</p> <p>Validation:</p> <p>These labels have been extensively validated and further information can be referenced in associated publications:<br><a href="https://doi.org/10.3390/rs13122364">https://doi.org/10.3390/rs13122364</a><br><a href="https://doi.org/10.3390/rs17071267">https://doi.org/10.3390/rs17071267</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

DownloadData used for the Bibliometric Analysis_ML in Pest Control

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo36/100

Supplementary data for "The mechanistic functional landscape of Retinitis Pigmentosa: an ML-driven approach to drug repurposing"

<p>Supplementary data for "The mechanistic functional landscape of Retinitis Pigmentosa: an ML-driven approach to drug repurposing"</p><p>&nbsp;</p><p>version: 10.5281/zenodo.10203479</p><ul><li>added missing file: "drug_actions_withSimplAction.csv"</li></ul>

opencc-by-nc-4.0May 2023View details →
zenodo36/100

Improving Generalization of ML-based IDS with Lifecycle-based Dataset, Auto-Learning Features, and Deep Learning

<p>The dataset and CNN model used in this paper:</p> <p>Improving Generalization of ML-based IDS with Lifecycle-based Dataset, Auto-Learning Features, and Deep Learning</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Wikipedia: wikipedia-ml (Malayalam)

Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://ml.wikipedia.org/

opencc-by-sa-4.0Aug 2024View details →
zenodo36/100

Geoclaw inputs and processed outputs used in building tsunami ML surrogates for nearshore and onshore approximation

<p>This dataset contains some input files needed for GeoClaw tsunami simulation and post-processed outputs used for training the nearshore and onshore surrogates discussed in the article - Advancing nearshore and onshore tsunami hazard approximation with machine learning surrogates available as preprint (https://doi.org/10.5194/nhess-2024-72) and project repo - https://github.com/naveenragur/tsunami-surrogates</p> <p>The GeoClaw simulation requires the following files:&nbsp;</p> <p><strong><a href="../api/records/10817116/draft/files/geoclaw_dtopo_files.tar.gz/content" target="_blank" rel="noopener noreferrer">geoclaw_dtopo_files.tar.gz</a></strong><strong>: </strong>input bathymetry and topography elevation in asc format from multiple sources of datasets.</p> <p><strong><a href="../api/records/10817116/draft/files/dtopo_his.tar.gz/content" target="_blank" rel="noopener noreferrer">dtopo_his.tar.gz</a></strong>: tsunami displacement(dtopo) input files in tt3 format for historic earthquake scenarios.</p> <p><strong><a href="../api/records/10817116/draft/files/dtopo_typeA.tar.gz/content" target="_blank" rel="noopener noreferrer">dtopo_typeA.tar.gz</a>:</strong> tsunami displacement(dtopo) input files in tt3 format for DOE type A earthquake scenarios.</p> <p><strong><a href="../api/records/10817116/draft/files/dtopo_typeB.tar.gz/content" target="_blank" rel="noopener noreferrer">dtopo_typeB.tar.gz</a>:</strong> tsunami displacement(dtopo) input files in tt3 format for DOE type A earthquake scenarios.</p> <p>The machine learning training and testing requires:</p> <p><strong><a href="../api/records/10817116/draft/files/procesed_tsunami_data.tar/content" target="_blank" rel="noopener noreferrer">procesed_tsunami_data.tar</a>:&nbsp;</strong>the post-processed outputs for the three test location i.e. (1) the waveforms( time series recorded for the water level at the offshore and the nearshore gauge locations and (2) the maximum inundation depth recorded at the fixed grids onshore.</p> <p><a href="https://zenodo.org/uploads/14902165"><strong>tsunami-surrogates-nhess-2024-72.tar.gz</strong></a><a href="https://zenodo.org/uploads/14902165"> </a>: Model code, scripts and notebooks used in the study</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Met Office UKCP Local CPM precipitation ML emulator dataset

<div> <div> <div> <h1>Met Office UKCP Local CPM precipitation ML emulator dataset</h1> <p>This is a collection of two datasets: one sourced from CPM data (bham64_ccpm-4x_12em_psl-sphum4th-temp4th-vort4th_pr.tar.gz) and one sourced from GCM data (bham64_gcm-4x_12em_psl-sphum4th-temp4th-vort4th_pr.tar.gz). Each dataset is made up of climate model variables extracted from the Met Office's storage system, combining many variables over many years. It consists of 3 NetCDF files (train.nc, test.nc and val.nc), a YML ds-config.yml file and a README (similar to this one but tailored to the source of the data). Code used to create the dataset can be found here: <a href="https://github.com/henryaddison/mlde-data">https://github.com/henryaddison/mlde-data</a> (specifically the james-submission tag).</p> <p>The YML file contains the configuration for the creation of the dataset, including the variables, scenario, ensemble members, spatial domain and resolution, and the scheme for splitting the data across the three subsets.</p> <p>Each NetCDF contains the same variables but split into different subsets (train, val and test) of the based on time dimension.</p> <p>Otherwise the NetCDF files have the sames dimensions and coordinates for ensemble_member, grid_longitude and grid_latitude.</p> <ul> <li>Spatial resolution: This has two parts - the resolution of the data and the grid resolution stored at in the file. For predictand variables this is 2.2km variables coarsened 4 times to 8.8km (this is the target grid). For predictor variables this is 2.2km variables conservatively regriddded to GCM 60km grid or variables from GCM (so already on 60km grid) then regrid (nearest neighbour) to the target grid of predictands. In the naming convention of resolution used in config files, 60km resolution is synonamous with the GCM grid and 2.2km resolution is synonamous with the CPM grid.</li> <li>Spatial domain: A 64x64 section of the 8.8km target grid covering England and Wales</li> <li>Time resolution: daily</li> <li>Time domain: 1st Dec 1980 to 30th Nov 2000; 1st Dec 2020 to 30th Nov 2040; 1st Dec 2060 to 30th Nov 2080. Uses a 360-day calendar.</li> <li>Scenario: RCP8.5</li> <li>Ensemble Members: 01, 04-13 &amp; 15 (these correspond to the 12 ensemble member runs from the CPM but don't carry intrinsic meaning).</li> <li>Split scheme: 70% training, 15% validation, 15% testing, split by choosing complete seasons at random, with an equal number of each season from each of the 3 time periods.</li> </ul> <p>&nbsp;</p> <h2>Predictor variables</h2> <ul> <li>psl (hPa) - mean sea level pressure</li> <li>temp850, temp700, temp500, temp250 - air temperature (K) at 850, 700, 500 and 250 hPa</li> <li>vorticity850, vorticity700, vorticity500, vorticity250 - relative vorticity (s^-1) at 850, 700, 500 and 250 hPa</li> <li>spechum850, spechum700, spechum500, spechum250 - specific humidity at 850, 700, 500 and 250 hPa</li> </ul> <h2>Predictand variable</h2> <ul> <li>target_pr - precipitation rate (mm/day)</li> </ul> <p>&nbsp;</p> <p>UPDATE 2025-03-27: Dataset tars are renamed to make it clearer their source (ccpm for coarsened CPM and gcm for GCM).</p> </div> </div> </div>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Environmental watering volumes (ML) used (held and planned environmental water) on the Murray-Darling Basin between 2007-08 and 2020-21

<p>An assessment of the volumes of environmental water released in the Murray-Darling Basin, Australia, by catchment, between 2007-08 and 2020-21</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

A Reconstructed Global Daily Seamless TROPOMI SIF dataset at a 0.05-degree Resolution (SDSIF) using the ML approach

<p>To enhance the spatial and temporal resolutions and continuities of the TROPOMI SIF, a global daily seamless SIF product at 0.05-degree resolution (namely, SDSIF) from May 2018 to December 2020 was generated based on the ML approach using TROPOMI SIF, MODIS reflectance, and ERA5 reanalysis datasets. This dataset has been validated with the original TROPOMI SIF and the long-term tower-based SIF from five flux sites, which verified the reliability of SDSIF and the advantages over original TROPOMI SIF.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Training Data for Decision Tree Practical in reading-ml-chemistry repo

<p>This is a dataset that can be used to train a decision tree to predict the band gap of a material. The data is associated with a notebook for running the practical and it can be found at https://github.com/keeeto/reading-ml-chemistry.</p> <p>&nbsp;</p> <p>The data originally comes from the Materials Project.</p> <p>&nbsp;</p> <p>A new muon spectroscopy dataset is added. It is from - Machine learning approach to muon spectroscopy analysis - https://iopscience.iop.org/article/10.1088/1361-648X/abe39e/meta</p>

opencc-by-4.0Jan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record