Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,075
datasets available to search
ShareScore release 0.7.1
Dataset results
1,075 results for “ML”
SPEA TESTBED TEST DATA FOR ML DEVELOPMENT
<p>The dataset files were generated at 3 different temperature with the improved ATE system developed in WP3 of the MET4FOF Project. </p> <p>The data are available for ML purposes and metrological investigation.</p>
Prior C - posterior_statistics_using_ml
<p>Training set data (Prior C) as used in </p> <p>Hansen. T.M. and Finlay, C.C., Use of machine learning to estimate properties of the posterior distribution in probabilistic inverse problems - an application to airborne EM data. JGR - Solid Earth, 20/10/2022.</p> <p>doi:10.1029/2022JB024703<br> </p>
Prior A - posterior_statistics_using_ml
<p>Training set data (Prior A) as used in </p> <p>Hansen. T.M. and Finlay, C.C., Use of machine learning to estimate properties of the posterior distribution in probabilistic inverse problems - an application to airborne EM data. JGR - Solid Earth, 20/10/2022.</p> <p>doi:10.1029/2022JB024703<br> </p>
Open MatSci ML Toolkit - DGL Graphs for OpenCatalyst IS2RE Task (OC20)
<p>Dataset accompanying the release of the <a href="https://github.com/IntelLabs/matsciml">Open MatSciML Toolkit</a>, an open source software for development graph neural networks on the OpenCatalyst project using the Deep Graph Library (DGL). </p> <p>For more details about the Open MatSci ML Toolkit, check the associated <a href="https://github.com/IntelLabs/matsciml">open-source repository</a> and <a href="https://arxiv.org/abs/2210.17484">paper</a>.</p> <p>Compressed files ~8GB with uncompressed file being ~80 GB.</p>
PYFOREST-ML-Predictions
<p>Results data from the Master's Capstone project PYFOREST. Predictions of deforestation under different policy scenarios are provided. See attached metadata for additional information, or refer to the <a href="https://github.com/cp-PYFOREST">GitHub project repository.</a> <br> </p>
TableS1 of ML-based predictive gut microbiome analysis for health assessment
<p>Complete list of species associated with the COVID and Control cohort (ANOVA F score > 1). Comparison with respect to the original set of species used by Gupta <em>et al.</em>, and ANOVA-F values are reported.</p>
Study of Secukinumab With 2 mL Pre-filled Syringes
ClinicalTrials.gov study NCT02748863. IPD Sharing: UNDECIDED. Countries: 11. Publications: 1.
Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.
Fire-D: Analysis and ML-Ready NASA-Centric Remote Sensing of Wildfire and Smoke
<p>Earth science remote sensing imagery is rich in structural and spectral information, making such data an ideal platform for benchmarking for a broad range of machine learning (ML) tasks, from pattern retrieval to physics-informed classification to anomaly detection to transfer learning. Nevertheless, the utility of Earth science remote sensing data remains largely unexplored by the broader ML community. Our goal is to bridge this gap and bring a rich variety of multisource multi-resolution Earth image data to a wider range of ML researchers who are non-experts in remote sensing, thereby increasing the utility and societal impact of such data products. In particular, motivated by the emerging wildfire crisis, we present radiometrically and geometrically calibrated radiance data from airborne and orbital instruments from the National Aeronautics and Space Administration (NASA), the National Oceanic and Atmospheric Administration (NOAA), and the Korean Meteorological Administration (KMA).</p> <p>Given the scarce occurrence of wildfires and complex spatio-temporal dependencies in radiance data, these datasets are especially well suited for benchmarking unsupervised and self-supervised learning tasks both on images and non-Euclidean objects. Our experiments on these datasets indicate that contrastive learning and transfer learning algorithms can capture the structures of views and scenes, map pixel space of multi-sensor imagery to a high-level embedding space for further downstream tasks, and facilitate more cohesive integration of the state-of-the-art ML approaches into wildfire risk analytics.</p> <p>All NASA-based observations are freely usable under the <a href="https://science.data.nasa.gov/license/">Creative Commons Zero License</a>.There are also no restrictions on the use of <a href="https://registry.opendata.aws/noaa-goes/">GOES Data</a>. <a href="https://registry.opendata.aws/noaa-gk2a-pds/">GK2A data</a> are also open data without any restrictions on its use.<br><br>For the Planet data, we cannot not share the Radiances, but all masks within this dataset are freely usable with no restrictions.</p> <p> </p> <p>Use:</p> <p>On the data input, input geometrically and radiometrically calibrated radiance data has been pulled from various NASA, NOAA, Planet, and KMA archives. For instruments that have multiple different spatial resolutions within their spectral bands (GOES and GK2A), all bands have been resampled to the lowest collective spatial resolution.</p> <p>Geometric and radiometric calibration has been done by the science data processing pipelines of the various missions, and would not need to be done by anyone else looking to curate the same data. Further information for each instrument can be found in each of the publicly available Level-1 algorithm theoretical basis documents (ATBDs)</p> <p>All input and label data have been put in GeoTiff format. Each band is in a separate raster band and each scene is in a separate GeoTiff file. Label files and input files are in separate tar files, labeled respectively, and the file names match for input and labels, with the exception of an additional .fire and .smoke in the respective label filenames and subfolders.<br><br>The <a href="https://www.earthdata.nasa.gov/about/esdis/esco/standards-practices/geotiff">GeoTiff</a> data format natively contains geolocation metadata internally, and can be interfaced with via C/C++/Python <a href="https://gdal.org/en/stable">GDAL</a> packages, or other python packages that wrap GDAL, like <a href="https://rasterio.readthedocs.io/en/stable/">rasterio</a> and <a href="https://corteva.github.io/rioxarray/stable/">rioxarray</a> . The documentation for <a href="https://nicks-personal-organization-2.gitbook.io/sit-fuse">SIT-FUSE</a> , the package with which the labels were generated, also has examples on how to read and interface with various data formats, including GeoTiffs. Lastly, this data can be interfaced with using Geographic Information Systems (GIS), like the free and open-source <a href="https://qgis.org/">QGIS</a>.</p> <p>An example of programmatic data access and usage can be found in the dataset's associated <a href="https://github.com/Fire-D-Dataset/FIRE-D">GitHub repository</a>. </p> <p>A working example using data from this repository for ML tasks is available <a href="https://drive.google.com/drive/folders/16aJO6LhrxJ3gsWoTU9BNN3hsb8W0refG?usp=sharing">here</a>.</p> <p>Timing information can be found in the file names, which all use the standard formats from the various instruments' L1B datasets.</p> <p>V2 includes additional GOES-18 radiance data and associated smoke and fire labels for the recent LA fires (Palisades and Eaton fires in January of 2025).</p> <p>V3 provides a reorganization of all data, and an inclusion of improved and additional data from airborne and satellite platforms in 2019, associated with this study: https://arxiv.org/pdf/2501.15343 . </p> <p>V4 provides additional AVIRIS-C Radiances and fixes the spatial range of the GOES-17 radiances to match that of the associated labels. The AVIRIS-C radiances are split across 5 tar files, ordered temporally - all associated labels are in a single tar file.</p> <p><br>Current fire coverage includes:</p> <ul> <li>2019: Williams Flats, Sheridan, Horsefly, and Mosquito (US)</li> <li>2022: Uljin Forest Fire (S. Korea; largest fire on record in S. Korea)</li> <li>2025: Palisades and Eaton Fires (US)</li> </ul> <p>Additional data for the 2025 Palisades and Eaton fires from the TEMPO instrument is currently being validated and will be released in a V4 shortly.</p> <p>Croissant file for dataset metadata specification is also included</p> <p>Validation:</p> <p>These labels have been extensively validated and further information can be referenced in associated publications:<br><a href="https://doi.org/10.3390/rs13122364">https://doi.org/10.3390/rs13122364</a><br><a href="https://doi.org/10.3390/rs17071267">https://doi.org/10.3390/rs17071267</a></p> <p> </p> <p> </p>
DownloadData used for the Bibliometric Analysis_ML in Pest Control
Open the record for dataset details and reuse information.
Supplementary data for "The mechanistic functional landscape of Retinitis Pigmentosa: an ML-driven approach to drug repurposing"
<p>Supplementary data for "The mechanistic functional landscape of Retinitis Pigmentosa: an ML-driven approach to drug repurposing"</p><p> </p><p>version: 10.5281/zenodo.10203479</p><ul><li>added missing file: "drug_actions_withSimplAction.csv"</li></ul>
Improving Generalization of ML-based IDS with Lifecycle-based Dataset, Auto-Learning Features, and Deep Learning
<p>The dataset and CNN model used in this paper:</p> <p>Improving Generalization of ML-based IDS with Lifecycle-based Dataset, Auto-Learning Features, and Deep Learning</p>
Wikipedia: wikipedia-ml (Malayalam)
Wikipedia is a multilingual, web-based, free-content encyclopedia project supported by the Wikimedia Foundation and based on a model of openly editable content. EOL harvests articles from wikipedia that are indexed as species or higher taxa.<p></p><p></p>https://ml.wikipedia.org/
Geoclaw inputs and processed outputs used in building tsunami ML surrogates for nearshore and onshore approximation
<p>This dataset contains some input files needed for GeoClaw tsunami simulation and post-processed outputs used for training the nearshore and onshore surrogates discussed in the article - Advancing nearshore and onshore tsunami hazard approximation with machine learning surrogates available as preprint (https://doi.org/10.5194/nhess-2024-72) and project repo - https://github.com/naveenragur/tsunami-surrogates</p> <p>The GeoClaw simulation requires the following files: </p> <p><strong><a href="../api/records/10817116/draft/files/geoclaw_dtopo_files.tar.gz/content" target="_blank" rel="noopener noreferrer">geoclaw_dtopo_files.tar.gz</a></strong><strong>: </strong>input bathymetry and topography elevation in asc format from multiple sources of datasets.</p> <p><strong><a href="../api/records/10817116/draft/files/dtopo_his.tar.gz/content" target="_blank" rel="noopener noreferrer">dtopo_his.tar.gz</a></strong>: tsunami displacement(dtopo) input files in tt3 format for historic earthquake scenarios.</p> <p><strong><a href="../api/records/10817116/draft/files/dtopo_typeA.tar.gz/content" target="_blank" rel="noopener noreferrer">dtopo_typeA.tar.gz</a>:</strong> tsunami displacement(dtopo) input files in tt3 format for DOE type A earthquake scenarios.</p> <p><strong><a href="../api/records/10817116/draft/files/dtopo_typeB.tar.gz/content" target="_blank" rel="noopener noreferrer">dtopo_typeB.tar.gz</a>:</strong> tsunami displacement(dtopo) input files in tt3 format for DOE type A earthquake scenarios.</p> <p>The machine learning training and testing requires:</p> <p><strong><a href="../api/records/10817116/draft/files/procesed_tsunami_data.tar/content" target="_blank" rel="noopener noreferrer">procesed_tsunami_data.tar</a>: </strong>the post-processed outputs for the three test location i.e. (1) the waveforms( time series recorded for the water level at the offshore and the nearshore gauge locations and (2) the maximum inundation depth recorded at the fixed grids onshore.</p> <p><a href="https://zenodo.org/uploads/14902165"><strong>tsunami-surrogates-nhess-2024-72.tar.gz</strong></a><a href="https://zenodo.org/uploads/14902165"> </a>: Model code, scripts and notebooks used in the study</p>
Met Office UKCP Local CPM precipitation ML emulator dataset
<div> <div> <div> <h1>Met Office UKCP Local CPM precipitation ML emulator dataset</h1> <p>This is a collection of two datasets: one sourced from CPM data (bham64_ccpm-4x_12em_psl-sphum4th-temp4th-vort4th_pr.tar.gz) and one sourced from GCM data (bham64_gcm-4x_12em_psl-sphum4th-temp4th-vort4th_pr.tar.gz). Each dataset is made up of climate model variables extracted from the Met Office's storage system, combining many variables over many years. It consists of 3 NetCDF files (train.nc, test.nc and val.nc), a YML ds-config.yml file and a README (similar to this one but tailored to the source of the data). Code used to create the dataset can be found here: <a href="https://github.com/henryaddison/mlde-data">https://github.com/henryaddison/mlde-data</a> (specifically the james-submission tag).</p> <p>The YML file contains the configuration for the creation of the dataset, including the variables, scenario, ensemble members, spatial domain and resolution, and the scheme for splitting the data across the three subsets.</p> <p>Each NetCDF contains the same variables but split into different subsets (train, val and test) of the based on time dimension.</p> <p>Otherwise the NetCDF files have the sames dimensions and coordinates for ensemble_member, grid_longitude and grid_latitude.</p> <ul> <li>Spatial resolution: This has two parts - the resolution of the data and the grid resolution stored at in the file. For predictand variables this is 2.2km variables coarsened 4 times to 8.8km (this is the target grid). For predictor variables this is 2.2km variables conservatively regriddded to GCM 60km grid or variables from GCM (so already on 60km grid) then regrid (nearest neighbour) to the target grid of predictands. In the naming convention of resolution used in config files, 60km resolution is synonamous with the GCM grid and 2.2km resolution is synonamous with the CPM grid.</li> <li>Spatial domain: A 64x64 section of the 8.8km target grid covering England and Wales</li> <li>Time resolution: daily</li> <li>Time domain: 1st Dec 1980 to 30th Nov 2000; 1st Dec 2020 to 30th Nov 2040; 1st Dec 2060 to 30th Nov 2080. Uses a 360-day calendar.</li> <li>Scenario: RCP8.5</li> <li>Ensemble Members: 01, 04-13 & 15 (these correspond to the 12 ensemble member runs from the CPM but don't carry intrinsic meaning).</li> <li>Split scheme: 70% training, 15% validation, 15% testing, split by choosing complete seasons at random, with an equal number of each season from each of the 3 time periods.</li> </ul> <p> </p> <h2>Predictor variables</h2> <ul> <li>psl (hPa) - mean sea level pressure</li> <li>temp850, temp700, temp500, temp250 - air temperature (K) at 850, 700, 500 and 250 hPa</li> <li>vorticity850, vorticity700, vorticity500, vorticity250 - relative vorticity (s^-1) at 850, 700, 500 and 250 hPa</li> <li>spechum850, spechum700, spechum500, spechum250 - specific humidity at 850, 700, 500 and 250 hPa</li> </ul> <h2>Predictand variable</h2> <ul> <li>target_pr - precipitation rate (mm/day)</li> </ul> <p> </p> <p>UPDATE 2025-03-27: Dataset tars are renamed to make it clearer their source (ccpm for coarsened CPM and gcm for GCM).</p> </div> </div> </div>
Environmental watering volumes (ML) used (held and planned environmental water) on the Murray-Darling Basin between 2007-08 and 2020-21
<p>An assessment of the volumes of environmental water released in the Murray-Darling Basin, Australia, by catchment, between 2007-08 and 2020-21</p>
A Reconstructed Global Daily Seamless TROPOMI SIF dataset at a 0.05-degree Resolution (SDSIF) using the ML approach
<p>To enhance the spatial and temporal resolutions and continuities of the TROPOMI SIF, a global daily seamless SIF product at 0.05-degree resolution (namely, SDSIF) from May 2018 to December 2020 was generated based on the ML approach using TROPOMI SIF, MODIS reflectance, and ERA5 reanalysis datasets. This dataset has been validated with the original TROPOMI SIF and the long-term tower-based SIF from five flux sites, which verified the reliability of SDSIF and the advantages over original TROPOMI SIF.</p>
Training Data for Decision Tree Practical in reading-ml-chemistry repo
<p>This is a dataset that can be used to train a decision tree to predict the band gap of a material. The data is associated with a notebook for running the practical and it can be found at https://github.com/keeeto/reading-ml-chemistry.</p> <p> </p> <p>The data originally comes from the Materials Project.</p> <p> </p> <p>A new muon spectroscopy dataset is added. It is from - Machine learning approach to muon spectroscopy analysis - https://iopscience.iop.org/article/10.1088/1361-648X/abe39e/meta</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.