Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

24

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

24 results for “ML modeling”

Learn how ShareScore rates datasets ↗
zenodo48/100

UDP Synthetic Dataset for training ML time series models

<p>The dataset available has been produced by the &quot;Next-Generation IoT solutions for the universal supply chain&quot; (iNGENIOUS) project&rsquo;s consortium under EC grant agreement 957216, &nbsp;made publicly available as part of the Horizon 2020 Open Research Data Pilot (<a href="https://www.openaire.eu/what-is-the-open-research-data-pilot">ORD pilot</a>).<br> The European Commission is not liable for any use that may be made of the information contained herein.</p> <p>The available dataset is in csv format and contains synthetic data of UDP packets received and sent by a single User Plane Function (UPF) covering a span of 6 weeks. The format of the datafile is:</p> <ul> <li>index</li> <li>timestamp&nbsp;</li> <li>UDP packets_rcvd - Total number of UDP packets received</li> <li>UDP packets sent - Total number of UDP packets sent</li> </ul> <p>The simulation was performed based on behavior of UPF and 5GC Network functions inferred from stress tests performed in the iNGENIOUS project&#39;s Automated Robots with Heterogeneous Networks Use Case, as well as patterns in urban mobility taken from available UE datasets [NCS+19].</p> <p>More information on the iNGENIOUS project can be found on the project&rsquo;s website: <a href="https://ingenious-iot.eu/">https://ingenious-iot.eu/</a></p> <p>[NCS+19] Noussan M, Carioni G, Sanvito FD, Colombo E. Urban Mobility Demand Profiles:<br> Time Series for Cars and Bike-Sharing Use as a Resource for Transport and Energy<br> Modeling. Data. 2019; 4(3):108. https://doi.org/10.3390/data4030108</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

ML-Enabled Systems Model Deployment and Monitoring: Status Quo and Problems

<p>Contained within this directory is the latest dataset utilized in the research titled 'ML-Enabled Systems Model Deployment and Monitoring: Status Quo and Problems'. We are providing a downloadable ZIP file that includes the survey questionnaire, the amassed data, and the Jupyter Notebooks utilized for our analytical process.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Datasets for benchmarking and ML modelling

<p><em><span>hydrogen-harm</span></em><span> data set of crystalline hydrogen configurations: energies at VMC and LRDMS level; purpose: benchmark for MLP; developed in the group of Michele Casula (CNRS) </span></p> <p><span><em>prot-hex</em> data set for protonated water hexamer: trajectories from classical molecular dynamics with nuclear forces at VMC level of theory; purpose: ML modelling; developed in the group of Michele Casula (CNRS)</span></p> <p><span><em>intexcit</em> data sets for a set of organic molecular complexes in lowest excited states: dispersion interaction energies, interaction energies, components of SAPT interaction energies at the CAS wavefunction level; purpose: benchmarking <em>ab initio</em> methods and density functional dispersion correction modelling; developed by Kasia Pernal (TUL) and Michal Hapka (University of Warsaw) </span></p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Dataset and structure database for an ML model to predict diffusivity in ZIF variants

<p>This dataset accompanies the publication titled &quot;Data Mining for Predicting Gas Diffusivity in Zeolitic-imidazolate Frameworks (ZIFs)&quot; (DOI:&nbsp;<a href="https://doi.org/10.1039/D2TA02624D">https://doi.org/10.1039/D2TA02624D</a>)</p> <p><a href="https://zenodo.org/api/files/b80f6d07-3bf4-484c-97ac-5d579fb0cc27/ESI_2_dataset.xlsx?versionId=dc4525d0-1c5c-478c-9bef-a156587ad69b">ESI_2_dataset.xlsx</a>: Descriptors for all ZIFs of the publication and simulations output, in the form of diffusivities of gas molecules (He up to iso-butane), in all ZIFs.</p> <p>ZIF_database.zip: ZIP file containing all ZIFs prepared by the authors (as discussed in the publication), through various units replacements, in the SOD topology, in .pdb&nbsp;format.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Index based dataset for training ML classification models

<p>This dataset contains 202122 rows of data containing 61 unique indices from different world urban areas.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Forecasting 24-hour-averaged PM2.5concentration in the Aburrá Valley using tree-based ML models, global forecasts, and satellite information: Dataset

<p>Data necessary for the training and evaluating the 24-hourly-averaged PM2.5 forecast over 19 stations within the Aburr&aacute; Valley, Colombia,&nbsp;is included here.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

ReqExp: BERT-based ML Model for Extracting Software Requirements

<p>Datasets that were used during experiments in ReqExp project.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Modeled waver data for ML

<p>The uploaded data are numerical model generated&nbsp;daily wave height, period,&nbsp;and wind frocings in&nbsp;the Chesapeake Bay.&nbsp;Data are used for the publication &quot;Machine Learning-based Wave Model with High Spatial Resolution in Chesapeake Bay&quot; submitted to the Journal for review.</p>

opencc-byJul 2023View details →
zenodo32/100

Replication Package for 'Analyzing the Evolution and Maintenance of ML Models on Hugging Face'

<p>Replication Package attached to the 'Analyzing the Evolution and Maintenance of ML Models on Hugging Face' article. Within the README and accompanying scripts, you will find detailed instructions to guide you through the analysis conducted in the article.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Replication Package for "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models"

<p>This repository contains the replication package for the paper titled "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models". The README file and accompanying scripts provide detailed instructions to facilitate the replication of the analysis.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

ML models for dengue prediction using Orange

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

Using LLM Models and Explainable ML to Analyse Biomarkers at Single Cell Level

<p>Single-cell RNA sequencing (scRNA-seq) technology has significantly advanced our understanding of the diversity of cells and how this diversity is implicated in diseases. Yet, translating these findings across various scRNA-seq datasets poses challenges due to technical variability and dataset-specific biases. To overcome this, we present a novel approach that employs both an LLM-based framework and explainable machine learning to facilitate generalization across single-cell datasets and identify gene signatures to capture disease-driven transcriptional changes. Our approach uses scBERT, which harnesses shared transcriptomic features among cell types to establish consistent cell-type annotations across multiple scRNA-seq datasets. Additionally, we employ a symbolic regression algorithm to pinpoint highly relevant yet minimally redundant models and features for inferring a cell type&rsquo;s disease state based on its transcriptomic profile. We ascertain the versatility of these cell-specific gene signatures across datasets, showcasing their resilience as molecular markers to pinpoint and characterize disease-associated cell types. Validation is carried out using four publicly available scRNA-seq datasets from both healthy individuals and those suffering from ulcerative colitis (UC). This demonstrates our approach&rsquo;s efficacy in bridging disparities specific to different datasets, fostering comparative analyses. Notably, the simplicity and symbolic nature of the retrieved gene signatures facilitate their interpretability, allowing us to elucidate underlying molecular disease mechanisms using these models.</p>

openother-openSep 2023View details →
zenodo28/100

Dataset for ML model

<p>File with data</p>

opencc-by-4.0Oct 2022View details →
zenodo28/100

Testdataset for Downscaling with different ML models

<p>This dataset consists of ERA5 and CERRA data for a training period (2014), a validation period (some months in 2017), and a testing period (some months in 2018) for training a UNET on colab.</p>

opencc-by-4.0Jun 2023View details →
ClinicalTrials.gov28/100

ML Decision Model for G-NEC Adjuvant Therapy

ClinicalTrials.gov study NCT06663852. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov28/100

Anthropometric and US-Guided Difficult Intubation Prediction With ML Models

ClinicalTrials.gov study NCT06904586. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo24/100

Data set for training a ML model to predict duration of MPI application phases (HPC system) - with previous phase info

<p>This is the data used to train a ML model predicting the duration of MPI application phases, in a HPC system.</p> <p>There are 10 different data sets corresponding to different HPC applications.</p> <p>These data sets&nbsp; contain information regarding the previous MPI call with same ID and type.</p>

openApr 2019View details →
zenodo24/100

Data set for training a ML model to predict duration of MPI application phases (HPC system) - without previous phases info

<p>This is the data used to train a ML model predicting the duration of MPI application phases, in a HPC system.</p> <p>There are 11 different data sets corresponding to different HPC applications.</p> <p>These data sets&nbsp; do not contain information regarding previous MPI calls</p>

openApr 2019View details →
zenodo24/100

A Dataset for Applying Machine Learning and Eddy Covariance Approaches to Model Mangrove Carbon Production (ML-MCP)

<p>The Mangrove Carbon Production (ML-MCP) dataset (daily time scale) encompasses comprehensive measurements of carbon production in mangrove ecosystems from four EC tower station in the USA and China, derived using advanced machine learning models and eddy covariance techniques. This dataset includes various variables such as carbon fluxes, environmental factors. By integrating machine learning algorithms, the dataset enhances the accuracy of carbon productivity estimations, facilitating better understanding and management of mangrove ecosystems' role in carbon sequestration and climate regulation.</p>

opencc-by-4.0Sep 2024View details →
zenodo24/100

Replication Package for 'Exploring the Carbon Footprint of Hugging Face's ML Models: A Repository Mining Study'

<p>Replication Package attached to the &#39;Exploring the Carbon Footprint of Hugging Face&#39;s ML Models: A Repository Mining Study&#39; article. Within the README and accompanying scripts, you will find detailed instructions to guide you through the analysis conducted in the article.</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record