Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo40/100

SENTIMENT ANALYSIS OF CUSTOMER FEEDBACK IN THE BANKING SECTOR: A COMPARATIVE STUDY OF MACHINE LEARNING MODELS

<p><span>This study investigates the application of sentiment analysis to customer feedback in the banking sector, utilizing natural language processing (NLP) techniques and machine learning models to classify customer sentiments into positive, neutral, and negative categories. Feedback was sourced from online platforms, including bank websites, social media, and third-party review sites. Data preprocessing steps, such as tokenization, stemming, and feature extraction using TF-IDF, were employed to prepare the text for analysis. Various machine learning algorithms, including Logistic Regression, Random Forest, Support Vector Machine (SVM), Long Short-Term Memory (LSTM), and Na&iuml;ve Bayes, were implemented and evaluated using metrics such as accuracy, precision, recall, and F1-score. The results show that LSTM outperformed all models with a 91% accuracy, followed closely by SVM at 89%. These findings demonstrate the potential of advanced machine learning techniques in accurately classifying sentiments and provide valuable insights into customer satisfaction and areas for improvement within the banking sector. Future work aims to further optimize models for better classification of neutral feedback and explore more advanced deep learning models, such as BERT.</span></p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

A Dataset of the Operating Station Heat Rate for 806 Indian Coal Plant Units using Machine Learning

<div> <div> <div> <div> <p>India aims to achieve net-zero emissions by 2070 and has set an ambitious target of 500 GW of renewable power generation capacity by 2030. Coal plants currently contribute to more than 60% of India&rsquo;s electricity generation in 2022. Upgrading and decarbonizing high-emission coal plants became a pressing energy issue. A key technical parameter for coal plants is the operating station heat rate (SHR), which represents the thermal efficiency of a coal plant. Yet, the operating SHR of Indian coal plants varies and is not comprehensively documented. This study extends from several existing databases and creates an SHR dataset for 806 Indian coal plant units using machine learning (ML), presenting the most comprehensive coverage to date. Additionally, it incorporates environmental factors such as water stress risk and coal prices as prediction features to improve accuracy. This dataset, easily downloadable from our visualization platform, could inform energy and environmental policies for India&rsquo;s coal power generation as the country transitions towards its renewable energy targets.</p> </div> </div> </div> </div>

opencc-by-4.0Mar 2024View details →
zenodo40/100

MTClass: Identification and annotation of multi-phenotype cis-eQTLs using machine learning

<p>This dataset contains the aggregated results from three iterations of MTClass. Brief descriptions of the file names are below:</p> <ul> <li><strong>multi-tissue.zip</strong>: Multi-tissue study containing the 9-tissue, 13 brain tissue, and 48-tissue results from MTClass, MultiPhen, and MANOVA</li> <li><strong>multi-exon.zip</strong>: Multi-exon study containing the multi-exon results from the 13 individual brain tissues (MTClass, MultiPhen, and MANOVA for each tissue)</li> <li><strong>2D_exon_tissue.zip</strong>: Multi-tissue/multi-exon combined eQTL study, done in 9 tissues using multi-layer perceptron. MTClass was only run once on this dataset due to the relatively higher computational burden.</li> <li><strong>PsychENCODE_isoQTL.zip</strong>: Multi-isoform eQTL study, done in human prefrontal cortex using PsychENCODE data.</li> <li><strong>OneK1K_scRNAseq.zip</strong>: Multi-cell-type eQTL study, done using scRNA-seq data from the OneK1K cohort.</li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Four spatial prediction datasets of susceptibility to gully erosion, comparing machine learning models, in the Piraí drainage basin, southeastern Brazil

<p>The data in this repository refer to the article published in the journal Land, entitled: Machine Learning Models for the Spatial Prediction of Gully Erosion Susceptibility in the Pira&iacute; Drainage Basin, Para&iacute;ba do Sul Middle Valley, Southeast Brazil.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Fingerprint Matrix Files for "Machine Learning-based Bioactivity Classification of Natural Products Using LC-MS/MS Metabolomics"

<p>These files are the necessary dataset to reproduce the observed machine learning metrics in the paper "Machine Learning-based Bioactivity Classification of Natural Products Using LC-MS/MS Metabolomics" in review at the Journal of Natural Products.&nbsp;</p> <ul> <li>Multiclassifier_23_Drug_Class_Train-Test_Fingerprint_Matrix.tsv is the accumulated positive training set for the 23 different classes demonstrated in the training and testing sets.</li> <li>Negative_Train-Test_Fingerprint_Matrix.tsv is the negatives training and testing examples derived from the RIKEN NP Depo which represent a diverse set of natural product compounds that serve as the counter points to the positive examples.</li> <li>GNPS_23_Drug_Class_Fingerprints_Matrix.tsv is the dataset of fingerprints generated from the publically available GNPS MSMS dataset. These training examples serve to confirm the ability of the machine learning model to generalize to experimental data.&nbsp;</li> <li>&nbsp;Negative_Train-Test_Fingerprint_Matrix.tsv is the dataset of negative training examples derived from the publically available spectra from the GNPS dataset. It is composed of nearly 2,800 random MSMS spectra to compose a diverse negative evaluation set.&nbsp;</li> <li>Random_GNPS_Fingerprints.tsv is the dataset of fingeprints of 9,443 random spectra from GNPS used to evaluate the false positive rate of each model.</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Students' Academic Performance in Case-Based Learning (CBL) based on Machine Learning Approach

<p>Paper title: Students&rsquo; Academic Performance in Case-Based Learning (CBL) based on Machine Learning Approach.</p> <p>This paper was registered in 2024 7<sup>th</sup> International Seminar on Research of Information Technology and Intelligent Systems (ISRITI)</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Linked collectors and determiners for: Life beneath the ice: jellyfish and ctenophores from the Ross Sea, Antarctica, with an image-based training set for machine learning.

Natural history specimen data linked to collectors and determiners held within, "Life beneath the ice: jellyfish and ctenophores from the Ross Sea, Antarctica, with an image-based training set for machine learning". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/f1ddda5a-ac85-46a3-955c-b75277c6a600">https://bionomia.net/dataset/f1ddda5a-ac85-46a3-955c-b75277c6a600</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/f1ddda5a-ac85-46a3-955c-b75277c6a600">https://gbif.org/dataset/f1ddda5a-ac85-46a3-955c-b75277c6a600</a>. Formatted as a Frictionless Data package.

opencc-zeroJan 2024View details →
zenodo40/100

Supporting data for "The Operando Nature of Isobutene in H–SSZ–13 Unraveled by Machine Learning Potentials Beyond DFT Accuracy"

<p>Supporting data for "The Operando Nature of Isobutene in H&ndash;SSZ&ndash;13 Unraveled by Machine Learning Potentials Beyond DFT Accuracy" by M. Bocus, S. Vandenhaute and V. Van Speybroeck.</p> <p>This dataset contains examples of input files, submission and analysis scripts, in addition to the molecular dynamics trajectories used to obtain the results reported in the main manuscript. A more detailed description of the dataset content is provided in the README files therein.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Gaussian16 data for "Dynamic electronic structure fluctuations in the de novo peptide ACC-dimer revealed by first-principles theory and machine learning"

<p>This is the Gaussian 16 input and corresponding output, which was used as input into the machine learning presented in the paper titled "Dynamic electronic structure fluctuations in the de novo peptide ACC-dimer revealed by first-principles theory and machine learning". This upload is required before submission of the paper.<br><br>The 1001 and 100 snapshots from different extractions are preserved in separated directories. Each snapshot directory <code>*_snapshot</code> has the initial GROMACS snapshot <code>test_*.pdb</code> , the geometry after truncating the solvation shell in various formats, the Gaussian16 input, qsub input and the output directory <code>*.1</code> with a JobID number assigned by qsub. The output directory has the standard output from Gaussian in a <code>.log</code> file and <code>grep</code>ed output from the <code>.fchk</code>&nbsp; file in <code>*.out</code> .</p>

openmit-licenseOct 2024View details →
zenodo40/100

TRANSFORMING CUSTOMER RETENTION IN FINTECH INDUSTRY THROUGH PREDICTIVE ANALYTICS AND MACHINE LEARNING

<p>In recent years, the fintech industry has experienced rapid growth, driven by technological advancements and evolving consumer expectations. Fintech companies offer innovative financial services, such as digital banking, investment platforms, and payment solutions, catering to the needs of a tech-savvy customer base. However, as competition intensifies, customer retention has emerged as a critical challenge for these companies. According to a study by Ransom (2021), acquiring a new customer can cost five times more than retaining an existing one, making it imperative for fintech organizations to focus on strategies that enhance customer loyalty. The financial technology (fintech) sector has experienced unprecedented growth in recent years, fundamentally transforming how individuals and businesses access and manage financial services. Characterized by the integration of technology with financial services, fintech encompasses a wide array of offerings, including digital banking, peer-to-peer lending, robo-advisory services, and payment processing. As of 2023, the global fintech market was valued at approximately $309 billion and is projected to reach around $1.5 trillion by 2030, according to a report by Fortune Business Insights. This remarkable growth is largely attributed to advancements in digital technology, increasing smartphone penetration, and a growing consumer preference for online financial solutions. Moreover, the COVID-19 pandemic accelerated the adoption of digital financial services, as consumers sought contactless transactions and remote banking options.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

INNOVATIVE MACHINE LEARNING APPROACHES TO FOSTER FINANCIAL INCLUSION IN MICROFINANCE

<p>This study examines the application of machine learning algorithms to enhance financial inclusion in microfinance, focusing on credit scoring, risk and fraud detection, and customer segmentation. We performed feature engineering and employed models such as Logistic Regression, Decision Trees, Random Forests, Gradient Boosting Machines (XGBoost and LightGBM), Support Vector Machines (SVM), Autoencoders, Isolation Forests, and K-means Clustering. LightGBM achieved the highest accuracy (89.6%) and AUC (0.92) in credit scoring, while Random Forests demonstrated strong performance in both loan approval (86.7% accuracy) and fraud detection (87.6% accuracy, AUC of 0.88). SVM also performed competitively, and unsupervised methods like Autoencoders and Isolation Forests showed potential for anomaly detection but required further refinement.K-means Clustering excelled in customer segmentation with a silhouette score of 0.72, enabling tailored services based on client demographics. Our findings highlight the significant impact of machine learning on improving credit scoring accuracy, reducing fraud risks, and enhancing customer service delivery in microfinance, thereby promoting financial inclusion for underserved populations. Ethical considerations and model interpretability are crucial, particularly for smaller institutions. This study advocates for the broader adoption of machine learning in the microfinance sector.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

CsSn(Cl/Br/I)3 Perovskite Alloy DFT Dataset for Machine Learning

<p>This upload contains density functional theory (DFT) calculations of CsSn(Cl/Br/I)3 perovskite alloy. The calculations were performed for a study, where the DFT data was used to train an energy predicting machine learning model for CsSn(Cl/Br/I)3. The code related to the study is available through GitLab (https://gitlab.com/cest-group/learnsolar-cssnclbri).</p> <p>The data is divided into four data sets. For each set, the atomic structure data with total energies and forces has been separated into an ASE (Atomic Simulation Environment) extended XYZ file. Additional information on the atomic structures (e.g. space groups) is provided in JSON format. The data sets are:</p> <p><strong>sp_train_set</strong><br>Single point DFT calculations of 16 000 algorithmically generated CsSn(Cl/Br/I)3 structures of four different space groups: Pm-3m, P4/mbm, I4/mcm, and Pnma. Lattice parameters and atomic positions are determined through Vegard's law, but random deviations have been added to the atom positions, tilting angles of the Sn coordination octahedra, cell volume, cell height-to-width ratio, and some lattice vector angles. Cl/Br/I configurations are randomized. This data set was used to fit an initial machine learning model. The atomic structures included were selected using a clustering algorithm to accelerate learning.</p> <p><strong>sp_test_set</strong><br>Single point DFT calculations of 2 600 atomic structures similar to sp_train_set. The Cl/Br/I compositions are uniformly represented, having two atomic structures per composition and space group. This data was used for testing the machine learning model.&nbsp;</p> <p><strong>al_data</strong><br>DFT relaxation structure snapshots from the active learning run that was performed to improve the machine learning model's structure relaxation accuracy. There are 4230 structure snapshots in total.</p> <p><strong>relax_test_set</strong><br>100 DFT relaxations used for testing the machine learning relaxation accuracy. There are 2881 structure snapshots in total. Both initial (relax_test_set_initial.xyz) and final (relax_test_set_relaxed.xyz) atomic geometries are included.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Machine learning based multi-scale remodelling code

<p>This instruction illustrates a machine learning-based multi-scale model to predict bone formation in tissue scaffolds. This code uses neural networks to predict bone formation in synthetic scaffolds. We are sorry that the code is a little bit messy as we are not good at coding. The code can be used to predict bone remodelling results in synthetic scaffolds in a multi-level way. Therefore, it enables to inversely identification of the bone remodelling related parameters from clinical data. In order to run the machine learning-based multiscale bone remodelling program. The following platforms are what you need:</p> <ol> <li>Abaqus v2016/v614</li> <li>Matlab R2020b</li> <li>An Abaqus plugin tool which can be downloaded from&nbsp;<a href="https://github.com/mhogg/pyvxray.git">https://github.com/mhogg/pyvxray.git</a></li> <li>Jupyter notebook with Python 3.</li> </ol> <p>&nbsp;</p> <p><strong>Here is a detailed description of the program</strong>.</p> <ol> <li><strong>demo_example and demo_pearson_opt.&nbsp;</strong>This document provides instructions on running a demo example and a demo example calculating Pearson&rsquo;s coefficient. The necessary functions for running the machine learning-based algorithm are located in the folder &ldquo;demo_example&rdquo;. In &quot;demo_example&quot;,&nbsp;Multiscale_boneRemodelling_ML is the main function to start the program. &ldquo;macro_umat&rdquo; is the user subroutine to pass the homogenized material properties to Abaqus. &ldquo;read_macro_1423&rdquo; is a post-process file to obtain necessary stress/strain data information. &ldquo;Sheep_macro_1423_C3D4.inp&rdquo; is the input file of the sheep tibia scaffold model. Before start, please change line 49 and line 58 of &ldquo;macro_umat&rdquo; file to your current work directory.&nbsp; In &ldquo;results&rdquo; file, the results are obtained by running the demo_program.&nbsp;In &ldquo;demo_pearson_opt&rdquo; file, there are inversely-identified virtual X-ray images and <em>in-vivo </em>X-ray images. The python code &quot;sheep3_6_9_8roi.ipynb&quot; can be found to calculate the Pearson&rsquo;s coefficient between the virtual X-ray images and in-vivo X-ray images. Jupyter notebook is required to run &ldquo;sheep3_6_9_8roi.ipynb&rdquo; to calculate the Pearson&rsquo;s coefficient in &quot;demo_pearson_opt&quot;.&nbsp;&ldquo;image_opt&rdquo; is the calculated Pearson&rsquo;s coefficient based on the inverse-identified case.</li> <li><strong>Micro_samples</strong> contains the files for the generation of micro RVE samples. &ldquo;microRVE.inp&rdquo; is the input file of the micro RVE for Abaqus. &ldquo;Micro_USDF1.for&rdquo; is the Fortran file that is used as a user subroutine in Abaqus to consider the adaptive bone density change in the bone regeneration area. &ldquo;read_microRVE&rdquo; is the post-process file to obtain strain and stress information.</li> <li><strong>Neural_network_training</strong> contains the files for the training of the 1<sup>st</sup> neural network for calculating the elastic tensor and the 2<sup>nd</sup> series of neural networks for calculating the unit SED components. In <strong>1<sup>st</sup>_neural_network</strong> file, a Matlab code is for the training of the neural network based on Matlab R2020b. The training data of the 1<sup>st</sup> round and the 2<sup>nd</sup> round are provided in the dataset file. In <strong>2<sup>nd</sup>_neural_network</strong> file, a Matlab code &ldquo;Train3D_SED&rdquo; is used for the training of 21 independent neural networks for predicting 21 unit SED components. &ldquo;dataset&rdquo; contains the training data of unit SED components.</li> <li><strong>ML_approach </strong>contains the files of the proposed machine learning-based approach and the trained neural networks. The Matlab code &ldquo;Multiscale_boneRemodelling _ML&rdquo; is the program for the proposed approach, which will call Abaqus for the finite element analysis at the macroscopic level. &ldquo;sheep_macro_1423_C3D4.inp&rdquo; is the input file of the in-silico model, including sheep tibia, bone fixation plate, screws and scaffold. &ldquo;macro_umat&rdquo; is the user subroutine file. &ldquo;read_macro_1423&rdquo; is used for post-process of FE data. &ldquo;Trained_model&rdquo; file includes all the trained neural networks.</li> <li><strong>Inverse_identification</strong> contains files for the inverse problem. &ldquo;image_analysis&rdquo; contains the 5000 virtual images generated by the proposed ML approach (&ldquo;postresults_imageupdate_5000). &ldquo;sheep3_6_9_8roi.ipynb&rdquo; is the python-based code for the images analysis and calculates the Pearson&rsquo;s coefficient. &ldquo;NN_pear&rdquo; is the trained neural network to output Pearson&rsquo;s coefficient. &ldquo;Pea_opt&rdquo; is the code to find the optimal bone remodelling parameters by multi-objective genetic algorithm. &ldquo;in-vivo X_ray&rdquo; includes the clinical X-ray images taken at different time points. &ldquo;inverse_identified_results&rdquo; includes the final inverse identified results.</li> </ol>

opencc-by-4.0Apr 2021View details →
zenodo40/100

Molecular dynamics simulation of tricaproin in gas phase using machine-learning potential ANI2x

<p>Tricaproin (Glycerol trihexanoate) is an example of a triglyceride molecule with very short alkyl tails attached to the glycerol moiety, and this deposit contains a 10 ns long simulation of tricaproin in a gas phase.</p> <p>The model chemistry (a.k.a. interaction potential or force field) is the machine-learning potential ANI2x implemented in python package torchANI, which has a close-to-DFT accuracy, yet low cost compared to DFT or other electronic structure theories.</p> <p>Molecular dynamics were run using ASE with Langevin integrator at a constant temperature of 310 K.</p> <p>The resulting trajectory was written every 1 ps (&quot;traj.h5&quot;, can be viewed in ASEgui), and gathered every 10 ps in a XTC format (&quot;traj.xtc&quot;, open in MDAnalysis, VMD, UnityMol, GROMACS tools ...).</p> <p>&nbsp;</p> <p>Detailed simulation settings are in the python script.</p> <p>&nbsp;</p> <p>This simulation was performed for the purpose of building a coarse grained Martini 3 model of this molecule.</p> <p>&nbsp;</p> <p>ANI2x https://doi.org/10.26434/chemrxiv.11819268.v1</p> <p>torchANI https://aiqm.github.io/torchani/index.html</p> <p>ASE https://wiki.fysik.dtu.dk/ase/tutorials/md/md.html#constant-temperature-md</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Data for publication of "Determining the sensitive parameters of WRF model for the prediction of tropical cyclones in the Bay of Bengal using Global Sensitivity Analysis and Machine Learning"

<p>The data are made available as part of the paper &quot;Determining the sensitive parameters of WRF model for the prediction of tropical cyclones in the Bay of Bengal using Global Sensitivity Analysis and Machine Learning&quot;, submitted to Geoscientific Model Development. This data set incorporates selected post-processed files needed to reproduce the results presented in the paper.</p> <p>The data contains six zip files, that are:</p> <ul> <li>Namelist.input files for the WRF model simulations of ten tropical cyclones</li> <li>WRF model simulation outputs using the default parameter values</li> <li>WRF model simulation outputs using the optimal parameter values (which give minimum RMSE value)</li> <li>IMDAA surface observations and IMERG precipitation data</li> <li>IMD observed tracks of ten tropical cyclones</li> <li>Ipython notebooks of sensitivity analysis and machine learning codes</li> </ul> <p>The remaining files are the ncl scripts that were used to obtain the figures. The ncl scripts used the data in the zip files.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Machine learning models, and training, validation and test datasets for: "Sequence determinants of human gene regulatory elements"

<p>This record contains the training, test and validation datasets used to train and evaluate the machine learning models in manuscript:</p> <p><strong>Sahu, Biswajyoti, et al. &quot;Sequence determinants of human gene regulatory elements.&quot; (2021).</strong></p> <p><br> This record contains also the final hyperparameter-optimized models for each training dataset/task combination described in the manuscript. The README-files provided with the record describe the datasets and models in more detail. The datasets deposited here are derived from the original raw data (GEO accession: GSE180158) as described in the Methods of the manuscript.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Datasets for the manuscript "In silico proof of principle of machine learning-based antibody design at unconstrained scale"

<p>The zip file contains dataset files for the manuscript &quot;In silico proof of principle of machine learning-based antibody design at unconstrained scale&quot;</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Data Models for Dataset Drift Controls in Machine Learning With Optical Images - Datasets

<p>This dataset accompanies the paper&nbsp;titled</p> <p><em>Data Models for Dataset Drift Controls in Machine Learning with Images</em><br> <br> that appeared in the Transactions on Machine Learning Research<br> <br> <a href="https://openreview.net/forum?id=I4IkGmgFJz">https://openreview.net/forum?id=I4IkGmgFJz</a><br> &nbsp;</p> <pre><code>@article{ oala2023data, title={Data Models for Dataset Drift Controls in Machine Learning With Optical Images}, author={Luis Oala and Marco Aversa and Gabriel Nobis and Kurt Willis and Yoan Neuenschwander and Mich{\`e}le Buck and Christian Matek and Jerome Extermann and Enrico Pomarico and Wojciech Samek and Roderick Murray-Smith and Christoph Clausen and Bruno Sanguinetti}, journal={Transactions on Machine Learning Research}, issn={2835-8856}, year={2023}, url={https://openreview.net/forum?id=I4IkGmgFJz}, note={} }</code></pre> <p>We make available two datasets.</p> <p><strong>Raw-Microscopy:</strong></p> <ul> <li><strong>940 raw bright-field microscopy images</strong> of human blood smear slides for leukocyte classification (microscopy/images/raw_scale100) with corresponding labels (microscopy/labels).</li> <li><strong>5,640 variations measured at six additional different intensities </strong>(microscopy/images/raw_scale001-raw_scale0075)</li> <li><strong>11,280 images of the raw sensor data processed through twelve different pipelines</strong> (microscopy/images/processed_views)</li> </ul> <p><strong>Raw-Drone:</strong></p> <ul> <li><strong>548 raw drone camera images for car segmentation</strong> (drone/images_tiles_256/raw_scale100) with corresponding binary segmentation mask (drone/masks_tiles_256). The images and the masks are cropped from 12 raw drone camera images (drone/images_full/raw_scale100) and 12 masks (drone/masks_full) of size 3648 by 5472.</li> <li><strong>3,288 variations measured at six additional different intensities</strong> (drone/images_tiles_256/raw_scale001-raw_scale075).</li> <li><strong>6,576 images of the raw sensor data processed through twelve different pipelines</strong> (drone/images_tiles_256/processed_views).</li> </ul> <p>Detailed datasheets for the two datasets can be found in the appendices of the TMLR paper.</p> <p>The code repository for this project can be found at&nbsp;<a href="https://github.com/aiaudit-org/raw2logit">https://github.com/aiaudit-org/raw2logit</a></p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Rapid identification of MRSA using mass spectrometry and machine learning from over 20000 clinical isolates

<p>Rapidly&nbsp;identifying&nbsp;methicillin-resistant&nbsp;Staphylococcus&nbsp;aureus&nbsp;(MRSA)&nbsp;with&nbsp;high&nbsp;integration&nbsp;in&nbsp;the&nbsp;current&nbsp;workflow&nbsp;is&nbsp;critical&nbsp;in&nbsp;clinical&nbsp;practices.&nbsp;We&nbsp;proposed&nbsp;a MALDI-TOF MS based&nbsp;machine&nbsp;learning&nbsp;model&nbsp;for rapidly MRSA prediction,&nbsp;the model&nbsp;was&nbsp;evaluated&nbsp;on&nbsp;a&nbsp;prospective&nbsp;test&nbsp;and&nbsp;four&nbsp;external&nbsp;clinical&nbsp;sites.&nbsp;On&nbsp;the&nbsp;dataset&nbsp;comprising&nbsp;20359&nbsp;clinical&nbsp;isolates,&nbsp;the&nbsp;area&nbsp;under&nbsp;the&nbsp;receiver&nbsp;operating&nbsp;curve&nbsp;of&nbsp;the&nbsp;classification&nbsp;model&nbsp;was&nbsp;0.78&ndash;0.88. Our MALDI&ndash;TOF MS-based ML model for the rapid&nbsp;MRSA&nbsp;identification can&nbsp;be&nbsp;easily&nbsp;integrated&nbsp;into&nbsp;the&nbsp;current&nbsp;clinical&nbsp;workflows&nbsp;and can further support&nbsp;physicians&nbsp;prescribe&nbsp;proper&nbsp;antibiotic&nbsp;treatments.&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

2021 UN Open GIS Challenge 1 - Training on Satellite Data Analysis and Machine Learning with QGIS (Satellite_QGIS)

<p>This dataset is part of the&nbsp;<a href="https://www.osgeo.org/foundation-news/2021-osgeo-un-committee-educational-challenge/?fbclid=IwAR0UvwkPO2pay7C0tJawb63eewjBGfeL9TIQpYUFccza9OIo6HAolmHXLWE">2021 UN Open GIS Challenge 1 - Training on Satellite Data Analysis and Machine Learning with QGIS (Satellite_QGIS)</a>,</p> <p>Exercise 1:&nbsp;Supervised Change Detection: Monitoring deglaciation in Huascaran, Peru.</p>

opencc-by-4.0Sep 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record