Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

15

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

15 results for “machine learning (ML)”

Learn how ShareScore rates datasets ↗
zenodo48/100

ML-TOMCAT V2.0: Machine-Learning-Based Satellite-Corrected Global Stratospheric Ozone Profile Dataset

<p>MLTOMCAT V2 is 46 years (1979-2024) of gap free ozone profile data sets that is created by correcting biases in a TOMCAT Chemical Transport Model (CTM) simulated ozone profiles. We use Random Forest regression model to correct model biases.&nbsp;</p> <p>Each file contain monthly mean zonal mean ozone profiles. There are 6 data files.</p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_vmr_V2.nc</a>&nbsp;contains ozone profiles on&nbsp;geometric height levels (1 to 60 km) &nbsp;in&nbsp;mixing ratio units, whereas&nbsp;&nbsp;<a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_nd_V2.nc</a>&nbsp;contains ozone profile in number density units.</p> <p>Similarly,&nbsp;</p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_vmr_V2.nc</a>&nbsp;contains ozone profiles on 43 MLS pressure levels&nbsp;&nbsp;(1000 to 0.1&nbsp;hPa) &nbsp;in&nbsp;mixing ratio units, whereas&nbsp;&nbsp;<a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_nd_V2.nc</a>&nbsp;contains ozone profile in number density units.</p> <p>Please note that data below 300 hPa (~8km) and 1 hPa (~50 km) should be used with caution.</p> <p>There are two straospheric column files</p> <p>ML-TOMCAT-SCO_120ppb_boundary_V2_197901-202412.nc and</p> <p>ML-TOMCAT-SCO_150ppb_boundary_V2_197901-202412.nc</p> <p>Stratospheric column files calculated using 120 ppb and 150 ppb as a chemical ozone boundaries.</p> <p>A manuscript describing MLTOMCAT would be published in EESD (Dhomse et al., 2021).</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

A map of global peatland extent created using machine learning (Peat-ML)

<p>Map of global peatland extent estimated by machine learning. The download includes both a netcdf file version and a GeoTIFF (as a zip archive)</p> <p>Abstract from associated paper:</p> <p>Peatlands store large amounts of soil carbon and freshwater, constituting an important component of the global carbon and<br> hydrologic cycles. Accurate information on the global extent and distribution of peatlands is presently lacking but is needed<br> by Earth System Models (ESMs) to simulate the effects of climate change on the global carbon and hydrologic balance. Here,<br> we present Peat-ML, a spatially continuous global map of peatland fractional coverage generated using machine learning<br> techniques suitable for use as a prescribed geophysical field in an ESM. Inputs to our statistical model follow drivers of<br> peatland formation and include spatially distributed climate, geomorphological and soil data, along with remotely-sensed<br> vegetation indices. Available maps of peatland fractional coverage for 14 relatively extensive regions were used along with<br> mapped ecoregions of non-peatland areas to train the statistical model. In addition to qualititative comparisons to other maps<br> in the literature, we estimated model error in two ways. The first estimate used the training data in a blocked leave-one-out<br> cross-validation strategy designed to minimize the influence of spatial autocorrelation. That approach yielded an average r<sup>2</sup><br> of 0.73 with a root mean squared error and mean bias error of 9.11% and -0.36%, respectively. Our second error estimate<br> was generated by comparing Peat-ML against a high-quality, extensively ground-truthed map generated by Ducks Unlimited<br> Canada for the Canadian Boreal Plains region. This comparison suggests our map to be of comparable quality to mapping<br> products generated through more traditional approaches, at least for boreal peatlands.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Brazilian Cohort for Predicting Cardiovascular Events Using Machine Learning (PRE-CARE ML project)

<p>The project PRE-CARE ML addresses the development and internal and external validation of predictive models for the assessment of risks of major adverse cardiovascular events.&nbsp;</p> <p>Global and local interpretability analyses of predictions were conducted towards improving model reliability and tailoring preventive interventions.&nbsp;</p> <p>The models were trained and validated in a retrospective cohort with the use of data from Hospital das Cl&iacute;nicas da Faculdade de Medicina de Ribeirao Preto, Brazil</p> <p>The International Classification of Diseases (ICD-10) classified patients with MACE (case group). (see Table I for the ICD-10 codes that defined MACE).</p> <p>Table 1: ICD-10 Codees for MACE definition</p> <table> <tbody> <tr> <td><strong>ICD-10</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>I20</td> <td>Angina pectoris</td> </tr> <tr> <td>I21</td> <td>Acute myocardial infarction</td> </tr> <tr> <td>I24</td> <td>Other acute ischaemic heart diseases</td> </tr> <tr> <td>I46</td> <td>Cardiac arrest</td> </tr> <tr> <td>I63</td> <td>Cerebral infarction</td> </tr> <tr> <td>I64</td> <td>Stroke, not specified as hemorrhage or infarction</td> </tr> <tr> <td>I71</td> <td>Aortic aneurysm and dissection</td> </tr> <tr> <td>I74</td> <td>Arterial embolism and thrombosis</td> </tr> </tbody> </table> <p>Only a patient&rsquo;s first MACE was considered, and all previous hospitalizations within a 5-year window were considered MACE. The control group (non-MACE) involved hospitalizations with no MACE and no death within a 5-year window (between 2017 and 2022). A sample of 6,000 MACE (labeled as 1) cases and 12,000 non-MACE (labeled as 0) ones was constructed with data from HCFMRP for training and internal validation purposes. Another balanced MIMIC IV sample of 8,000 MACE cases and 8,000 nonMACE cases was used for external validation.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Bio-ML: Machine Learning-Friendly Biomedical Datasets for Equivalence and Subsumption Ontology Matching

<p>&nbsp;</p> <blockquote> <p><strong>This version is used in the Bio-ML track of the OAEI 2024; the only change compared to the OAEI 2023 is the deletion of certain training subsumption mappings.</strong></p> </blockquote> <p>&nbsp;</p> <h3><strong>Overview</strong></h3> <p>The purpose of these datasets is to support&nbsp;<em>equivalence</em> and <em>subsumption</em> ontology matching.</p> <p>There are five ontology pairs extracted from MONDO and UMLS:</p> <table> <tbody> <tr> <td>Source</td> <td>Task</td> <td>Category</td> <td>#SrcCls</td> <td>#TgtCls</td> <td>#Ref (equiv)</td> <td>#Ref (subs)</td> </tr> <tr> <td>Mondo</td> <td>OMIM-ORDO</td> <td>Disease</td> <td>9,648</td> <td>9,275</td> <td>3,721</td> <td>103</td> </tr> <tr> <td>Mondo</td> <td>NCIT-DOID</td> <td>Disease</td> <td>15,762</td> <td>8,465</td> <td>4,686</td> <td>3,338 (-1)</td> </tr> <tr> <td>UMLS</td> <td>SNOMED-FMA</td> <td>Body</td> <td>34,418</td> <td>88,955</td> <td>7,256</td> <td>5,453 (-53)</td> </tr> <tr> <td>UMLS</td> <td>SNOMED-NCIT</td> <td>Pharm</td> <td>29,500</td> <td>22,136</td> <td>5,803</td> <td>4,224 (-1)</td> </tr> <tr> <td>UMLS</td> <td>SNOMED-NCIT</td> <td>Neoplas</td> <td>22,971</td> <td>20,247</td> <td>3,804</td> <td>213</td> </tr> </tbody> </table> <p>The "-" numbers reflect the changes due to lthe deletion of certain training subsumption mappings.</p> <p>The main track is available at "bio-ml", where each pair is associated with a task folder, containing the source and target ontologies, reference equivalence mappings (in "refs_equiv"), reference subsumption mappings ("refs_subs").&nbsp;</p> <p>The special sub-track is available at "bio-llm", where each pair is associated with a task folder, containing the source and target ontologies, and the test candidate mappings.&nbsp;</p> <p>&nbsp;</p> <h3><strong>Citation</strong></h3> <p><strong>Bio-ML (Main Track)</strong></p> <pre>```<br>@inproceedings{he2022machine, title={Machine learning-friendly biomedical datasets for equivalence and subsumption ontology matching}, author={He, Yuan and Chen, Jiaoyan and Dong, Hang and Jim{\'e}nez-Ruiz, Ernesto and Hadian, Ali and Horrocks, Ian}, booktitle={International Semantic Web Conference}, pages={575--591}, year={2022}, organization={Springer} }<br>```</pre> <p><strong>Bio-LLM (Sub-track)</strong></p> <pre>```<br>@article{he2023exploring, title={Exploring large language models for ontology alignment}, author={He, Yuan and Chen, Jiaoyan and Dong, Hang and Horrocks, Ian}, journal={arXiv preprint arXiv:2309.07172}, year={2023} }<br>```</pre> <p>&nbsp;</p> <h3><strong>Important Links</strong></h3> <ul> <li>See detailed documentation at:&nbsp;<a href="https://krr-oxford.github.io/DeepOnto/bio-ml">https://krr-oxford.github.io/DeepOnto/bio-ml</a>.</li> <li>See the OAEI Bio-ML track at:&nbsp;<a href="https://www.cs.ox.ac.uk/isg/projects/ConCur/oaei/">https://www.cs.ox.ac.uk/isg/projects/ConCur/oaei/</a></li> <li>See our resource paper for the original Bio-ML at&nbsp;<a href="https://arxiv.org/abs/2205.03447">arxiv</a>&nbsp;or <a href="https://link.springer.com/chapter/10.1007/978-3-031-19433-7_33">springer</a>&nbsp;(accepted at&nbsp;<em>ISWC-2022</em> and nominated as the <em>best resource paper candidate</em>). See our poster paper for the Bio-LLM sub-track at&nbsp;<a href="https://arxiv.org/abs/2309.07172">arxiv </a>(accepted at <em>ISWC-2023 Posters &amp; Demos</em>).</li> </ul> <p>&nbsp;</p> <h3><strong>Changelog</strong></h3> <p>The only change in this version compared to the OAEI 2023 is the deletion of certain training subsumption mappings that can be directly exploited through deductive reasoning.</p>

opencc-by-4.0Jul 2022View details →
ClinicalTrials.gov36/100

Clinical Performance Evaluation of the Artificial Intelligence (AI)/ Machine Learning (ML) Technologies Utilized by the Origin Medical EXAM ASSISTANT

ClinicalTrials.gov study NCT06952439. IPD Sharing: NO. Countries: 1. Publications: 15.

closedIPD-NOFeb 2026View details →
zenodo32/100

Dataset for "The State of the ML-universe: 10 Years of Artificial Intelligence & Machine Learning Software Development on GitHub"

<p>Supplementary data to &quot;The State of the ML-universe: 10 Years of Artificial Intelligence &amp; Machine Learning Software Development on GitHub&quot; accepted for publication at MSR 2020.</p> <p>The data included in this package were used to conduct analyses to characterize the AI &amp; ML software development community hosted on GitHub. Please read the paper for a full understanding of what data was collected and how it was used.</p> <p>Questions and comments can be directed to Danielle Gonzalez dng2551@rit.edu</p>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Satellite and Celestial Data for Machine Learning (SCD-ML)

<p>These datasets contains information about Kosmos 2514 satellite, and can be used for machine learning. The data comes from the <a href="https://www.juntadeandalucia.es/institutodeestadisticaycartografia/" target="_blank" rel="noopener">Institute of Statistics and Cartography of Andalusia (IECA)</a>. Then there is a detailed explanation:</p> <ul> <li><em><strong>satellite_data:</strong></em> This dataset includes ephemerides &nbsp;(precise satellite position and velocity data)&nbsp; about the satellite</li> <li><em><strong>sgdp4_celestial_data</strong></em>: This dataset contains ephemerides for the Kosmos 2514 satellite, positions of various celestial bodies in the solar system, and SGDP4 predictions for the satellite's position.</li> <li><em><strong>sequential_data_smj</strong></em>: This dataset includes sequences of 10 positions for the Sun, Moon, and Jupiter.</li> <li><em><strong>sequential_data_svmmj</strong></em>:&nbsp;Similar to the previous dataset, this one contains sequences of 10 positions for the Sun, Venus, Moon, Mars, and Jupiter. Each sequence also spans from the initial position to the SGDP4 predicted position of the satellite.</li> </ul> <p>The last two datasets, <em><strong>sequential_data_smj </strong></em>and <em><strong>sequential_data_svmmj</strong></em>, only provide sequences that are linked to the corresponding rows in&nbsp;<em><strong>sgdp4_celestial_data</strong></em>. They detail 10 positions of celestial bodies over the period between the initial position and the SGDP4 predicted position of the satellite.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

VacSol-ML(ESKAPE) Machine Learning DataSet

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →
zenodo24/100

A Dataset for Applying Machine Learning and Eddy Covariance Approaches to Model Mangrove Carbon Production (ML-MCP)

<p>The Mangrove Carbon Production (ML-MCP) dataset (daily time scale) encompasses comprehensive measurements of carbon production in mangrove ecosystems from four EC tower station in the USA and China, derived using advanced machine learning models and eddy covariance techniques. This dataset includes various variables such as carbon fluxes, environmental factors. By integrating machine learning algorithms, the dataset enhances the accuracy of carbon productivity estimations, facilitating better understanding and management of mangrove ecosystems' role in carbon sequestration and climate regulation.</p>

opencc-by-4.0Sep 2024View details →
ClinicalTrials.gov24/100

Feasibility and Utility of Artificial Intelligence (AI) / Machine Learning (ML) - Driven Advanced Intraoperative Visualization and Identification of Critical Anatomic Structures and Procedural Phases

ClinicalTrials.gov study NCT05775133. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov20/100

MUSCLE-ML: Multimodal Integration of Muscle Strength, Structure by Machine Learning for Precision Rehabilitation After ACL Injury

ClinicalTrials.gov study NCT07284771. IPD Sharing: Not stated. Countries: 0. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
geo16/100

Application of machine learning (ML) / deep learning (DL) using multiple epigenetic features reveals H3K27Ac as driver of gene expression prediction across patients with glioblastoma

GEO Series GSE296948. Homo sapiens. 11 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Expression profiling by high throughput sequencing.

openGEO-OpenJul 2025View details →
geo16/100

Application of machine learning (ML) / deep learning (DL) using multiple epigenetic features reveals H3K27Ac as driver of gene expression prediction across patients with glioblastoma [ATAC-seq]

GEO Series GSE296947. Homo sapiens. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenJul 2025View details →
geo16/100

Application of machine learning (ML) / deep learning (DL) using multiple epigenetic features reveals H3K27Ac as driver of gene expression prediction across patients with glioblastoma [RNA-Seq]

GEO Series GSE296945. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2025View details →
geo16/100

Application of machine learning (ML) / deep learning (DL) using multiple epigenetic features reveals H3K27Ac as driver of gene expression prediction across patients with glioblastoma [ChIP-Seq]

GEO Series GSE296944. Homo sapiens. 7 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenJul 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record