Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
geo16/100

An Optimized Hepatocyte-like Cells Differentiation Protocol Using Machine Learning Assisted Approach and Multi-omics Analysis [WGS]

GEO Series GSE246556. Homo sapiens. 6 samples. Type: Genome variation profiling by high throughput sequencing.

openGEO-OpenDec 2025View details →
geo16/100

The brain tumor classifications based on machine learning models trained with older version of methylation microarray chip compatible with the new EPIC v2 illumina’s chip.

GEO Series GSE229715. Homo sapiens. 32 samples. Type: Methylation profiling by genome tiling array.

openGEO-OpenSep 2023View details →
geo16/100

Machine learning-assisted classification of responders and non-responders to bortezomib treatment regimens PAD and VCD using new experimental dataset of 58 multiple myeloma RNA sequencing profiles

GEO Series GSE159426. Homo sapiens. 58 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2021View details →
geo16/100

Application of machine learning (ML) / deep learning (DL) using multiple epigenetic features reveals H3K27Ac as driver of gene expression prediction across patients with glioblastoma [RNA-Seq]

GEO Series GSE296945. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2025View details →
geo16/100

Functional Optimization of Designer Cardiac Organoids Enabled by Machine Learning Techniques

GEO Series GSE267438. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2024View details →
zenodo16/100

The microglia cytotoxicity dataset of Manuscript ci-2023-017235, entitled "Prediction and interpretation microglia cytotoxicity by machine learning"

<p>&nbsp;The microglia cytotoxicity dataset of manuscript ci-2023-017235, entitled "Prediction and interpretation microglia cytotoxicity by machine learning", in Journal of Chemical Information and Modeling. Including all molecular SMILES and their cytotoxicity.</p>

restrictedcc-by-4.0Oct 2023View details →
zenodo16/100

Diversity-aware Fairness Testing of Machine Learning Models through Hashing-based Sampling

<h4>VBT-X</h4><p>The results of vbt-x.</p><p>```</p><p>asdasd</p><p>```</p>

restrictedcc-by-4.0Nov 2023View details →
zenodo16/100

Data and Script for: Olive Trees Health and Yield Prediction through EO data and Machine Learning (OLEA-PRED) / EO AFRICA – Research and Development Facility

<p>This database was collected under the "Olive Trees Health and Yield Prediction through EO data and Machine Learning" project, funded by the European Space Agency in the framework of the "EO AFRICA R&amp;D Facility".</p> <p>- Shapefiles.rar contains shp files for Orchads boundaries and tree locations.</p> <p>- Field_data contains data collected in the field (Chllorophyl, Yield and Soil ).</p> <p>- SCRIPTS.rar contains all scripts and extracted data used for prediction for Settat orchad.</p> <p>- FBS images.rar contains UAV and M6 images for Fkih Ben Saleh orchad.</p>

embargoedcc-by-4.0Apr 2024View details →
zenodo16/100

Dataset para Estimação de Torque de Motores de Combustão Interna Flex-Fuel Utilizando Machine Learning

<p>Dataset foi projetado para auxiliar na modelagem e estima&ccedil;&atilde;o do torque de motores de combust&atilde;o interna do tipo Flex Fuel utilizando t&eacute;cnicas de Machine Learning.&nbsp;</p>

restrictedcc-by-4.0Nov 2024View details →
zenodo16/100

Machine Learning Prediction Models for Mitral Valve Repairability and Mitral Regurgitation Recurrence in Patients Undergoing Surgical Mitral Valve Repair

<p>This record contains raw data related to the article &quot;Machine Learning Prediction Models for Mitral Valve Repairability and Mitral Regurgitation Recurrence in Patients Undergoing Surgical Mitral Valve Repair&quot;</p> <p>Abstract: Background: Mitral valve regurgitation (MR) is the most common valvular heart disease and current variables associated with MR recurrence are still controversial. We aim to develop a machine learning-based prognostic model to predict causes of mitral valve (MV) repair failure and MR recurrence. Methods: 1000 patients who underwent MV repair at our institution between 2008 and 2018 were enrolled. Patients were followed longitudinally for up to three years. Clinical and echocardiographic data were included in the analysis. Endpoints were MV repair surgical failure with consequent MV replacement or moderate/severe MR (&gt;2+) recurrence at one-month and moderate/severe MR recurrence after three years. Results: 817 patients (DS1) had an echocardiographic examination at one-month while 295 (DS2) also had one at three years. Data were randomly divided into training (DS1: n = 654; DS2: n = 206) and validation (DS1: n = 164; DS2 n = 89) cohorts. For intra-operative or early MV repair failure assessment, the best area under the curve (AUC) was 0.75 and the complexity of mitral valve prolapse was the main predictor. In predicting moderate/severe recurrent MR at three years, the best AUC was 0.92 and residual MR at six months was the most<br> important predictor. Conclusions: Machine learning algorithms may improve prognosis after MV repair procedure, thus improving indications for correct candidate selection for MV surgical repair.</p>

restrictedJan 2022View details →
zenodo16/100

Machine Learning solution for machining quality prediction using acoustic emissions, accelerometers and current data

<p>This outcome contains the dataset used to train tool wear &amp; quality prediction algorithms in a milling context. The data model is composed of three data sources (acoustics emission, accelerometers &amp; currents). The acoustics emissions and accelerometers are recorded on an external machine which needs a post-synchronization. The currents are recorded by the system that directly monitors the machine.</p> <p><em>Acoustics emission (file:&nbsp;toolwear_2020_ae.zip):</em>&nbsp;In a previous data acquisition (not published), the acoustic emission was recorded using three sensors (on the raw part, on the axis closest to the part &amp; outside the machine), but in this dataset, it is only composed of one sensor positioned inside the machine. The sampling rate is 200kHz.</p> <p><em>Accelerometers (file: toolwear_2020_acc.zip):</em>&nbsp;The accelerometer dataset is composed of nine different sensors with an acquisition frequency of 20kHz. One signal is dedicated to the synchronization between other data sources (acoustics emission &amp; currents). Five signals are installed on the spindle axis, two (XY directions) on the top of the spindle and three (XYZ directions) on the bottom of the axis. The last three (XYZ directions) are located on the axis nearest to the part.</p> <p><em>Currents (file:&nbsp;toolwear_2020_axes.zip):</em>&nbsp;The machine is composed of five axes and one spindle. For each motor, the current is acquired and stored into the monitoring system of the CNC. The frequency data acquisition is 1kHz.</p>

restrictedApr 2022View details →
zenodo16/100

Dataset related to article "An individualized algorithm to predict mortality in COVID-19 pneumonia: a machine learning based study "

<p>This record contains raw data related to article &ldquo;An individualized algorithm to predict mortality in COVID-19 pneumonia: a machine learning based study&quot;</p> <p>Abstract:</p> <p><strong>Introduction: </strong> Identifying SARS-CoV-2 patients at higher risk of mortality is crucial in the management of a pandemic. Artificial intelligence techniques allow one to analyze large amounts of data to find hidden patterns. We aimed to develop and validate a mortality score at admission for COVID-19 based on high-level machine learning.</p> <p><strong>Material and methods: </strong> We conducted a retrospective cohort study on hospitalized adult COVID-19 patients between March and December 2020. The primary outcome was in-hospital mortality. A machine learning approach based on vital parameters, laboratory values and demographic features was applied to develop different models. Then, a feature importance analysis was performed to reduce the number of variables included in the model, to develop a risk score with good overall performance, that was finally evaluated in terms of discrimination and calibration capabilities. All results underwent cross-validation.</p> <p><strong>Results: </strong> 1,135 consecutive patients (median age 70 years, 64% male) were enrolled, 48 patients were excluded, and the cohort was randomly divided into training (760) and test (327) groups. During hospitalization, 251 (22%) patients died. After feature selection, the best performing classifier was random forest (AUC 0.88 &plusmn;0.03). Based on the relative importance of each variable, a pragmatic score was developed, showing good performances (AUC 0.85 &plusmn;0.025), and three levels were defined that correlated well with in-hospital mortality.</p> <p><strong>Conclusions: </strong> Machine learning techniques were applied in order to develop an accurate in-hospital mortality risk score for COVID-19 based on ten variables. The application of the proposed score has utility in clinical settings to guide the management and prognostication of COVID-19 patients.</p>

restrictedOct 2022View details →
zenodo16/100

Nature of ‎Metal-Support Interaction Discovered by Interpretable Machine ‎‎Learning

<p>Data for the figures and MD simulation movies in the publication "Nature of Metal-Support Interaction Discovered by Interpretable Machine &lrm;&lrm;Learning".</p>

restrictedcc-by-4.0Apr 2024View details →
zenodo16/100

Detection and tracking of turbulent structures in the edge of tokamak plasmas using machine learning

<p>This repository contains the data used in the study entitled "Detection and tracking of turbulent structures in the edge of tokamak plasmas: use of an ultra-fast camera and comparison of machine learning and Kalman filter methods". The data was collected using an ultra-fast camera on the COMPASS device.</p> <p><strong><em>This repository includes:</em></strong></p> <p><strong>Fast passive imaging data:</strong> High-frequency captures of turbulent structures present in the edge of tokamak plasmas.<br><strong>Labels :</strong> Annotation of turbulent structures for training and validation of detection and tracking algorithms.<br><strong>Detection and tracking results using YOLO:</strong> Outputs from the YOLO (You Only Look Once) model applied to turbulent data.<br><br><strong><em>Objective</em></strong><br>This dataset is intended to provide a complete and reproducible set for research into the detection and tracking of turbulent structures in tokamak plasmas. It can be used to compare the performance of machine learning methods with traditional techniques, and to encourage further research in this crucial area for plasma physics and nuclear fusion.</p> <p><br>Researchers are encouraged to use these data to:</p> <p>Replicate the results of the original study.<br>Develop and test new techniques for detecting and tracking turbulent structures.Compare the performance of machine learning algorithms with conventional methods.</p> <p><br>References Please cite the repository as follows: S. Chouchene et al., (2024). Detection and tracking of turbulent structures in the edge of tokamak plasmas using machine learning. Zenodo. https://doi.org/10.5281/zenodo.12608068</p>

restrictedcc-by-4.0Jun 2024View details →
zenodo16/100

Code for Quantitative Assessment of Factors Influencing Heat Vulnerability in Residential Areas using Machine Learning and UAV Data

<div> <div>Author: Ja Woon Gu</div> <div>Email: umseakind2@kwater.or.kr</div> <div>Date: 2024-07-22</div> <div>Version: 1.0</div> <br> <div>Description:</div> <div>This script compares multiple regression models using GridSearchCV and RepeatedKFold cross-validation.</div> <div>The script identifies the best performing model based on the mean cross-validation score (neg_mean_squared_error).</div> <div>It performs residual analysis including normality, homoscedasticity, and autocorrelation checks.</div> <br> <div>Models included:</div> <div>- DecisionTreeRegressor</div> <div>- ExtraTreesRegressor</div> <div>- AdaBoostRegressor</div> <div>- XGBRegressor</div> <div>- LGBMRegressor</div> <div>- CatBoostRegressor</div> <div>- RandomForestRegressor</div> <div>- GradientBoostingRegressor</div> <br> <div>The best model's feature importances are calculated and plotted.</div> <div>Residuals are analyzed using Shapiro-Wilk, Levene's, and Durbin-Watson tests.</div> </div>

restrictedcc-by-4.0Jul 2024View details →
zenodo16/100

Machine learning for identifying risk of death in patients with severe fever with thrombocytopenia syndrome

<p>Raw data of the manuscript</p>

restrictedcc-by-4.0Aug 2024View details →
zenodo16/100

Mapping Soil Organic Carbon in the World's Largest Arid Mangrove Forest (Indus Delta, Pakistan): A Multi-Sensor Remote Sensing and Machine Learning Approach

<p><span>Mangrove forests play a crucial role in carbon sequestration, especially in arid regions where their ability to store carbon in soil is vital for mitigating climate change. The Indus Delta in Pakistan, the world&rsquo;s largest arid mangrove forest system, lacks spatially explicit data on Soil Organic Carbon (SOC) despite its importance for conservation and carbon budgeting. This study aims to establish a baseline SOC map 2020 at 10 m spatial resolution using Sentinel-1 (Synthetic Aperture Radar) and Sentinel-2 (MultiSpectral Instrument) satellite imagery, integrated with in-situ soil sampling. SOC predictions were made using a Classification and Regression Tree (CART) machine learning model within the Google Earth Engine platform, leveraging 40 predictor variables, including spectral bands and derived indices. A total of 53 topsoil (0-10 cm) samples were collected in February 2020 across the Indus Delta, and SOC was analyzed using the Walkley-Black method. The results showed an average SOC value of 65.88 Mg C ha</span><span>⁻</span><span>&sup1; with substantial spatial variability, ranging from 15.06 Mg C ha</span><span>⁻</span><span>&sup1; to 138.03 Mg C ha</span><span>⁻</span><span>&sup1; with a total of 0.91 Pg C. The CART model demonstrated high accuracy, with an R&sup2; of 0.95 and an RMSE of 9.18 Mg C ha</span><span>⁻</span><span>&sup1;. However, the region faces challenges such as seawater intrusion and salinity, which threaten its ability to sequester carbon. With the first high-resolution SOC map for the Indus Delta, this study provides valuable insights for ecosystem management, conservation planning, and carbon budgeting. These findings of this study have the potential to significantly influence initiatives like REDD+ and Blue Carbon projects, which aim to enhance carbon sequestration while addressing the ecological challenges facing Pakistan&rsquo;s mangroves</span></p>

restrictedcc-by-4.0Sep 2024View details →
zenodo16/100

Refractive index determination of dynamic droplets in a flow by analyzing light scattering signals with a machine learning approach

<p>This container includes the measurement data, python script and weights of trained machine learning model associated with the scientific work, which will be presented in 2025 at the <em><strong>Turbulence, Heat and Mass Transfer 11</strong> </em>conference in Tokyo.</p> <p><strong>Title:</strong> Refractive Index Determination of Dynamic Droplets in Flow by Analyzing Light Scattering Signals with a Machine Learning Approach &nbsp;<br><strong>Authors:</strong> W. Schaefer<br><strong>Affiliation:</strong> ai-quanton GmbH, Dr.-Werner-Freyberg-Str. 7, 69514 Laudenbach, Germany &nbsp;<br><strong>Contact:</strong> info@ai-quanton.com&nbsp;</p> <p>The following data files are provided:</p> <ul> <li><strong>Dataset_40_4ch1234.rar (unpacked: Dataset_40_4ch1234.pth)</strong></li> <li><strong>M1_SegmentsTHR40.csv</strong></li> <li><strong>SegmentsTHR40.rar (unpacked: M1_SegmentsTHR40.csv ... M55_SegmentsTHR40.csv)</strong></li> <li><strong>Model_weights_4ch1234.pth</strong></li> </ul> <p>&nbsp;</p> <p><strong>Dataset_40_4ch1234.pth</strong> is a file, containing a ready-to-use dataset of 4-channel signals prepared for use in Python scripts.</p> <p><strong>M1_SegmentsTHR40.csv </strong>is an example of a file used for storing and loading light scattering signals of individual droplets with corresponding additional data. The meaning of each column is:</p> <p>'MID' &ndash; measurement ID</p> <p>'FID' &ndash; frame ID</p> <p>'SID' &ndash; signal ID</p> <p>'CID' &ndash; channel ID</p> <p>'NOP' &ndash; number of parts</p> <p>'PNM' &ndash; part number</p> <p>'TCH' &ndash; trigger channel</p> <p>'TLE' &ndash; trigger level</p> <p>'TID' &ndash; trigger ID</p> <p>'CON' &ndash; label used for training</p> <p><strong>SegmentsTHR40.rar</strong> is an archived folder containing .csv files, the same format as M1_SegmentsTHR40.csv.</p> <p><strong>Model_weights_4ch1234.pth </strong>contains weights for a model trained on data from all 4 channels.</p> <p>&nbsp;</p> <p><strong>External files:</strong></p> <p>The correcponding repository to this dataset is published on Azure Dev Ops: <a href="https://dev.azure.com/ai-quanton/PBa202">https://dev.azure.com/ai-quanton/PBa202</a><br>This repository contains the Python script developed for a neural network that determines the refractive index of single droplets by analyzing light scattering signals generated as they pass through a Gaussian beam.&nbsp;</p> <p>The script is designed to build and test a machine learning model capable of accurately predicting refractive indices from light scattering data in dynamic spray environments.</p>

restrictedcc-by-4.0Oct 2024View details →
zenodo16/100

Spectra from "Contribution of MALDI-TOF mass spectrometry and Machine Learning including Deep Learning techniques for the detection of virulence factors of Clostridioides difficile strains"

<p><strong>This database includes spectra from 201 <em>C. difficile&nbsp;</em> (CD) strains :</strong></p> <ul> <li>50 non-toxigenic strains (tcdA- tcdB-) (designated ToxA-B-) belonging to 19 different PR,</li> <li>151 toxigenic strains harbouring toxins A and B genes (ToxA+B+). Among the 151 ToxA+B+ strains, 46 corresponding to 8 different PR also harboured the binary toxin genes (ToxA+B+CDT+) and 105 (23 different PR) did not (ToxA+B+CDT-).</li> <li>Among the 46 ToxA+B+CDT+ strains, 22 belonged to the Hv strains i.e. PR 027 (n=13), PR 176 (n=5) and PR 181 (n=4) strains (ToxA+B+CDT+Hv) (Table S1).&nbsp;</li> </ul> <p><strong>Sample preparation.</strong> Each isolate stored at &minus;80&deg;C (Microbank; Pro-Lab Diagnostics) was thawed and cultivated on Columbia Blood Agar (CBA, bioM&eacute;rieux) incubated in anaerobic atmosphere at 37&deg;C for 48 hours. A subculture was performed in the same conditions. A chemical protein extraction was then carried out. Briefly, a single colony was suspended in 200 &micro;l water and vortexed. After adding 900 &micro;l ethanol, samples were vortexed and centrifuged at 13,000 &times; g for 2 minutes. The supernatant was removed, and the remaining ethanol was evaporated at room temperature. Next, 25 &micro;l of 70% formic acid was added and mixed with the pellet, then 25 &micro;l of acetonitrile was added. After centrifugation at 13,000 &times; g for 2 minutes, the supernatant was ready for analysis. <strong>Eight deposits were performed for each isolate.</strong> The dried spots were coated with 1 &micro;l of &alpha;-cyano-4-hydroxycinnamic acid (a-HCCA) in 50% acetonitrile-2.5% trifluoroacetic acid and <strong>each spot was analysed three times by MALDI-TOF MS</strong>.</p> <p><br> <strong>MALDI-TOF MS acquisition and analysis.</strong> Mass spectra were acquired using a Microflex LT instrument (Bruker Daltonics). The standard parameters of the CE-IVD method recommended by the manufacturer were used. This instrument was equipped with an N2 laser (377 nm) using the following parameters: mass range, 2,000 to 20,000 Da; ion source 1, 20 kV; ion source 2, 18.15 kV; lens, 6 kV; pulsed ion extraction, 150 ns; laser frequency, 20 Hz. A manual external calibration standard (Bacterial Test Standard; Bruker Daltonics) was used for calibration. Data acquisition was performed using FlexControl (version 3.0; Bruker Daltonics).</p> <p><br> <strong>A total of 4659 spectra were produced.&nbsp;</strong></p> <p><strong>Fore more details: please contact alexandre.godmer@aphp.fr</strong><br> &nbsp;</p>

restrictedOct 2023View details →
geo16/100

Constructing a rat model of stress cardiomyopathy and screening for diagnostic markers using machine learning algorithms

GEO Series GSE223385. Rattus norvegicus. 20 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record