Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo12/100

Machine Learning and MVPA

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Nov 2024View details →
zenodo12/100

Evaluation of Football Fan Comments on X Platform with Sen-timent Analysis and Machine Learning: The Case of Turkey

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Dec 2024View details →
zenodo12/100

[OUTDATED] Data set [ref. paper "Predictive modeling of drivers' brake reaction time through machine learning methods"]

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Dec 2024View details →
zenodo12/100

Prediction of illness remission in patients with Obsessive-Compulsive Disorder with supervised machine learning

<p>Prediction of illness remission in patients with Obsessive-Compulsive<br> Disorder with supervised machine learning</p> <p>&nbsp;</p> <p>Introduction: The course of OCD differs widely among OCD patients, varying from chronic symptoms to full<br> remission. No tools for individual prediction of OCD remission are currently available. This study aimed to<br> develop a machine learning algorithm to predict OCD remission after two years, using solely predictors easily<br> accessible in the daily clinical routine.<br> Methods: Subjects were recruited in a longitudinal multi-center study (NOCDA). Gradient boosted decision<br> trees were used as supervised machine learning technique. The training of the algorithm was performed with 227<br> predictors and 213 observations collected in a single clinical center. Hyper-parameter optimization was performed<br> with cross-validation and a Bayesian optimization strategy. The predictive performance of the algorithm<br> was subsequently tested in an independent sample of 215 observations collected in five different centers.<br> Between-center differences were investigated with a bootstrap resampling approach.<br> Results: The average predictive performance of the algorithm in the test centers resulted in an AUROC of<br> 0.7820, a sensitivity of 73.42%, and a specificity of 71.45%. Results also showed a significant between-center<br> variation in the predictive performance. The most important predictors resulted related to OCD severity, OCD<br> chronic course, use of psychotropic medications, and better global functioning.<br> Limitations: All recruiting centers followed the same assessment protocol and are in The Netherlands. Moreover,<br> the sample of the data recruited in some of the test centers was limited in size.<br> Discussion: The algorithm demonstrated a moderate average predictive performance, and future studies will<br> focus on increasing the stability of the predictive performance across clinical settings.</p>

restrictedJan 2022View details →
zenodo12/100

Inactive-enriched machine-learning models exploiting patent data improve structure-based virtual screening for PDL1 dimerizers

<p>The 12 VS scenarios considered in this study employing six training-test data partitions<strong> </strong>(A-F). All training sets employ the same set of 371 actives (WO2015160641A2), but differ on the considered set of inactives and hence are uniquely identified by the latter (either TrueInactives, DeepCoys, RandomDecoys or ActivesOnly). Likewise, all test sets employ the same 297 actives (WO201503820A1), none of them also included in the training set, but different sets of inactives (TrueInactives or DeepCoys).&nbsp;</p> <p>&nbsp;</p> <table align="center"> <caption>Table 1. Six virtual screening scenarios corresponding to six pairs of training-test data for each type of SFs (classification or regression)</caption> <thead> <tr> <th scope="col">Partition ID</th> <th scope="col">Training set</th> <th scope="col">Test set</th> <th scope="col">Type</th> </tr> </thead> <tbody> <tr> <td>A</td> <td>DeepCoys</td> <td>TrueInactives</td> <td>Classification</td> </tr> <tr> <td>B</td> <td>RandomDecoys</td> <td>TrueInactives</td> <td>Classification</td> </tr> <tr> <td>C</td> <td>ActivesOnly</td> <td>TrueInactives</td> <td>Classification</td> </tr> <tr> <td>D</td> <td>TrueInactives</td> <td>DeepCoys</td> <td>Classification</td> </tr> <tr> <td>E</td> <td>RandomDecoys</td> <td>DeepCoys</td> <td>Classification</td> </tr> <tr> <td>F</td> <td>ActivesOnly</td> <td>DeepCoys</td> <td>Classification</td> </tr> <tr> <td>A</td> <td>DeepCoys</td> <td>TrueInactives</td> <td>Regression</td> </tr> <tr> <td>B</td> <td>RandomDecoys</td> <td>TrueInactives</td> <td>Regression</td> </tr> <tr> <td>C</td> <td>ActivesOnly</td> <td>TrueInactives</td> <td>Regression</td> </tr> <tr> <td>D</td> <td>TrueInactives</td> <td>DeepCoys</td> <td>Regression</td> </tr> <tr> <td>E</td> <td>RandomDecoys</td> <td>DeepCoys</td> <td>Regression</td> </tr> <tr> <td>F</td> <td>ActivesOnly</td> <td>DeepCoys</td> <td>Regression</td> </tr> </tbody> </table> <p>&nbsp;</p>

restrictedFeb 2022View details →
zenodo12/100

Machine Learning to Predict In-Hospital Mortality in COVID-19 Patients Using Computed Tomography-Derived Pulmonary and Vascular Features

<p>Dataset from&nbsp;Schiaffino S, Codari M, Cozzi A, Albano D, Al&igrave; M, Arioli R, Avola E, Bn&agrave; C, Cariati M, Carriero S, Cressoni M, Danna PSC, Della Pepa G, Di Leo G, Dolci F, Falaschi Z, Flor N, Fo&agrave; RA, Gitto S, Leati G, Magni V, Malavazos AE, Mauri G, Messina C, Monfardini L, Pasch&egrave; A, Pesapane F, Sconfienza LM, Secchi F, Segalini E, Spinazzola A, Tombini V, Tresoldi S, Vanzulli A, Vicentin I, Zagaria D, Fleischmann D, Sardanelli F. Machine Learning to Predict In-Hospital Mortality in COVID-19 Patients Using Computed Tomography-Derived Pulmonary and Vascular Features. J Pers Med. 2021 Jun 3;11(6):501. doi: 10.3390/jpm11060501. PMID: 34204911; PMCID: PMC8230339.</p> <p>Abstract</p> <p>Pulmonary parenchymal and vascular damage are frequently reported in COVID-19 patients and can be assessed with unenhanced chest computed tomography (CT), widely used as a triaging exam. Integrating clinical data, chest CT features, and CT-derived vascular metrics, we aimed to build a predictive model of in-hospital mortality using univariate analysis (Mann-Whitney&nbsp;<em>U</em>&nbsp;test) and machine learning models (support vectors machines (SVM) and multilayer perceptrons (MLP)). Patients with RT-PCR-confirmed SARS-CoV-2 infection and unenhanced chest CT performed on emergency department admission were included after retrieving their outcome (discharge or death), with an 85/15% training/test dataset split. Out of 897 patients, the 229 (26%) patients who died during hospitalization had higher median pulmonary artery diameter (29.0 mm) than patients who survived (27.0 mm,&nbsp;<em>p</em>&nbsp;&lt; 0.001) and higher median ascending aortic diameter (36.6 mm versus 34.0 mm,&nbsp;<em>p</em>&nbsp;&lt; 0.001). SVM and MLP best models considered the same ten input features, yielding a 0.747 (precision 0.522, recall 0.800) and 0.844 (precision 0.680, recall 0.567) area under the curve, respectively. In this model integrating clinical and radiological data, pulmonary artery diameter was the third most important predictor after age and parenchymal involvement extent, contributing to reliable in-hospital mortality prediction, highlighting the value of vascular metrics in improving patient stratification.</p>

restrictedFeb 2022View details →
zenodo12/100

Ozone formation sensitivity study using machine learning coupled with the reactivity of volatile organic compound species

<p>Ozone formation sensitivity study using machine learning coupled with the reactivity of volatile organic compound species</p>

restrictedMar 2022View details →
zenodo12/100

Data for the manuscript "Classification of Stream, Hyperconcentrated, and Debris Flow Using Dimensional Analysis and Machine Learning"

<p>The excel file&nbsp;contains&nbsp;hydrological and sediment data.&nbsp;Also included are dimensional analysis data in our dataset.<br> The rar file contains the codes and data for SVM classification work.</p>

restrictedJul 2022View details →
zenodo12/100

Optimizing Radar-Based Rainfall Estimators Using Machine Learning Modles

<p>Weather radar research has produced numerous radar-based rainfall estimators based on climate, rainfall intensity, a variety of ground-truthing instruments and sensors (e.g., rain gauges, disdrometers), and techniques. Although each research direction gives improvement, their collective application in an operational sense still yields uncertainty in rainfall estimation at different times. This study aims to explore the concept of implementing Machine Learning (ML) models in choosing the optimal radar-based rainfall estimator from a group of estimators at each bin of a radar scan.&nbsp;&nbsp;</p> <p>The Canadian King City C-Band radar was used with a GEONOR T-200B rain gauge, a total of 263 sample points, to establish a group of polarimetric-based rainfall estimators (R(Z), R(Z, ZDR), R(KDP)). The estimators were used to train three ML models, namely Decision Tree, Random Forest, and Gradient Boost, to choose the optimal rainfall estimators based on radar variables (Z, ZDR, KDP). Data from the Canadian Exeter C-Band radar and a Texas Electronics TE525 tipping bucket gauge at a different location were used to verify the ML models and compare their results to the classic Marshall-Gunn (1952) Z-R relation and the composite estimator produced by Bringi et al. (2011). The results show promising results for the ML models, specifically the Gradient Boost model. These encouraging results need to be further explored with more sample points to further refine the ML mod</p>

restrictedNov 2022View details →
zenodo12/100

Revisiting Machine Learning based Test Case Prioritization for Continuous Integration

<p>Aborted version</p>

restrictedAug 2022View details →
zenodo12/100

Database of antimicrobial resistant bacterial taxa for supervised machine learning

<p>It is a database of antimicrobial resistant bacterial taxa for supervised machine learning.</p>

restrictedcc-by-4.0Jun 2024View details →
zenodo12/100

Machine Learning Potential for Electrochemical Interfaces with Hybrid Representation of Dielectric Response

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Jul 2024View details →
zenodo12/100

Raw data used to build models in SIMON Automated Machine Learning

<p>Here you can find all data and all information regarding each generated dataset.<br> For each dataset there are 4 files:</p> <p>json_info : This file contains, number of features with their names and number of subjects that are available for the same dataset<br> data_testing: data frame with data used to test trained model<br> data_training: data frame with data used to train models<br> results: direct unfiltered data from database</p> <p><br> Files are written in feather format.</p> <p><a href="https://gist.github.com/LogIN-/00d7628e0850f843ba84a678fac0a103">Here is an example</a> of data structure for each file in repository</p> <p>&nbsp;</p>

restrictedJul 2018View details →
zenodo12/100

DEMoS: an Italian emotional speech corpus. Elicitation methods, machine learning, and perception

<p>DEMoS (Database of Elicited Mood in Speech), is a corpus of induced emotional speech in Italian. DEMoS encompasses 9,365 emotional and 332 neutral samples produced by 68 native speakers (23 females, 45 males) in seven emotional states: the 'big six' anger, sadness, happiness, fear, surprise, disgust, and the secondary emotion guilt. To get more realistic productions, instead of acted speech, DEMoS contains emotional speech elicited by combinations of Mood Induction Procedures (MIP). Three elicitation methods are presented, made up by the combination of at least three MIPs, and considering six different MIPs in total. To select samples 'typical' of each emotion, evaluation strategies based on self- and external assessment were applied. The selected part of the corpus encompasses 1,564 prototypical samples produced by 59 speakers (21 females, 38 male). DEMoS has been published in the Journal Language, Resousrces, and Evalaution.</p> <p>&nbsp;</p> <p>Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Maximilian Schmitt, and Bj&ouml;rn Schuller (2019), <em>DEMoS: An Italian emotional speech corpus. Elicitation methods, machine learning, and perception</em>, Language, Resources, and Evaluation, Feb 2019. <a href="http://em.rdcu.be/wf/click?upn=lMZy1lernSJ7apc5DgYM8eCoqdGxOfRWEudjYRrxU-2BI-3D_Ru5N6PJ4ngeR7K-2Fncs2CW1jGAzl4dMvrVh77-2BVH-2B9g5urNss1KItQNXvWL1jiHKvcYDtUVs2c78DX20PMDTauCGehGiQvHdgrAknGggtu7pHINBqVKjp16-2BTn63kNrm22m52e-2FPV-2FidpRe8A-2FplLxPMV-2FjTR-2FLLIK8Wqe7u0-2BLSZ9w-2BWYtrAXRYn2lvPcjGTP1La8yiTxBuJKbHJpnNeFb6LmBIiNMmGRSZPIY0leXhyj4k07rx5cETF6n34aIQHP-2FwcafanNMN4BoA9QKhXGgFxvRgZQidsQ-2BCDbbTBL0PPjM3CgitSGk66qut9E3pd">https://rdcu.be/bn7oI</a></p> <p>&nbsp;</p> <p><strong>How to access DEMoS</strong></p> <p>To get access to the dataset, please send the signed End User License Agreement (EULA) when making the request. The EULA <strong>must be signed by somebody from a university holding a permanent position</strong>, typically a full professor. Note that requests without an EULA appropriately filled out, as well as those performed from a non-institutional e-mail address, will be automatically rejected. Please download the EULA from the following link:</p> <p>https://drive.google.com/file/d/1v6GaCVyNcib5v802t2uXHYOioqIkoBQ-/view?usp=share_link</p>

restrictedFeb 2019View details →
zenodo12/100

Supplementary Data and Models of Melt-based Thermo-barometer for paper "'No Free Lunch' in Tabular Geochemical Data: An example of Shallow versus Deep Machine Learning Algorithms for Geothermobarometry"

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Jun 2024View details →
zenodo12/100

A country-wide assessment of Iran's land subsidence susceptibility using satellite-based InSAR and machine learning

<p><em><strong>landsubsidance_ANN.tiff</strong></em>:&nbsp; Land-subsidence susceptibility potential maps using ANN algorithm.</p> <p><em><strong>landsubsidance_ghdm.tiff</strong></em>: Land-subsidence susceptibility potential maps using GMDH algorithm.</p>

restrictedJun 2021View details →
zenodo12/100

Predicting the skin sensitization potential of small molecules with machine learning models trained on biologically meaningful descriptors

<p>In recent years a number of machine learning models for the prediction of the skin sensitization potential of small organic molecules have been reported and become available. These models generally perform well within their applicability domains but, as a result of the use of molecular fingerprints and other non-intuitive descriptors, the interpretability of the existing models is clearly limited. The aim of this work is to develop a strategy to replace the non-intuitive features by predicted outcomes of bioassays. We show that such replacement is indeed possible and that as few as ten interpretable, predicted bioactivities are sufficient to reach competitive performance. On a holdout data set of 257 compounds, the best model (&quot;Skin Doctor CP:Bio&quot;) obtained an efficiency of 0.82 and an MCC of 0.52 (at the significance level of 0.20). Skin Doctor CP:Bio is available from the authors free of charge for academic research. The modeling strategies explored in this work are easily transferable and could be adopted for the development of more interpretable machine learning models for the prediction of the bioactivity and toxicity of small organic compounds.</p> <p>The corresponding research article has been published in <em>Pharmaceuticals</em> <strong>2021</strong>, <em>14</em>(8), 790, DOI: <a href="https://doi.org/10.3390/ph14080790">https://doi.org/10.3390/ph14080790</a></p>

restrictedJul 2021View details →
zenodo12/100

Dataset for "Classification of Stream, Hyperconcentrated, and Debris Flow Using Dimensional Analysis and Machine Learning"

<p>Du J. et al., (2022). Dataset for &quot;Classification of Stream, Hyperconcentrated, and Debris Flow Using Dimensional Analysis and Machine Learning&quot;, Water Resources Research</p> <p>Table S1:&nbsp;Hydrologic Paramteres and Dimensionless Numbers of Debris Flows</p> <p>Table S2:&nbsp;Hydrologic Paramteres and Dimensionless Numbers of Hyperconcentrated Flows</p> <p>Table S3:&nbsp;Hydrologic Paramteres and Dimensionless Numbers of Stream Flows</p> <p>Table S4:&nbsp;Hydrologic Paramteres and Dimensionless Numbers of Lahars</p>

restrictedNov 2022View details →
zenodo12/100

Machine Learning approach to Classification of Resting-State EEG Microstates in stroke survivors

<p>Dataset -&nbsp;Machine Learning approach to Classification of Resting-State EEG Microstates in stroke survivors</p> <p>https://docs.google.com/spreadsheets/d/1MeEx9ysEC_tWyohqEmvc8Ey5mc5Z9NV2/edit#gid=160681293</p>

restrictedJan 2023View details →
zenodo12/100

Texture analysis and machine learning to predict water T2 and fat fraction from non-quantitative MRI of thigh muscles in Facioscapulohumeral muscular dystrophy

<p><strong>Introduction</strong>. This database includes the radiomic features used as covariates to train machine learning algorithms in the paper &ldquo;&nbsp; Texture analysis and machine learning to predict water T2 and fat fraction from non-quantitative MRI of thigh muscles in Facioscapulohumeral muscular dystrophy&rdquo;.</p> <p><strong>Purpose</strong>. Quantitative MRI (qMRI) plays a crucial role for assessing disease progression and treatment response in neuromuscular disorders, but the required MRI sequences are not routinely available in every center. The aim of this study was to predict qMRI values of water T2 (wT2) and fat fraction (FF) from conventional MRI, using texture analysis and machine learning.</p> <p><strong>Method</strong>. Fourteen patients affected by Facioscapulohumeral muscular dystrophy were imaged at both thighs using conventional and quantitative MR sequences. Muscle FF and wT2 were calculated for each muscle of the thighs. Forty-seven texture features were extracted for each muscle on the images obtained with conventional MRI. Multiple machine learning regressors were trained to predict qMRI values from the texture analysis dataset.</p> <p><strong>Results</strong>. Eight machine learning methods (linear, ridge and lasso regression, tree, random forest (RF), generalized additive model (GAM), k-nearest-neighbor (kNN) and support vector machine (SVM) provided mean absolute errors ranging from 0.110 to 0.133 for FF and 0.068 to 0.115 for wT2. The most accurate methods were RF, SVM and kNN to predict FF, and tree, RF and kNN to predict wT2.</p> <p><strong>Conclusion</strong>. This study demonstrates that it is possible to estimate with good accuracy qMRI parameters starting from texture analysis of conventional MRI.</p>

restrictedJan 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record