Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: Hybrid machine learning approach to zero-inflated data improves accuracy of dengue prediction

Open the record for dataset details and reuse information.

publicDec 2025View details →
dryad36/100

Using machine learning to investigate premolar ecomorphology in anthropoid primates

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad36/100

A machine learning based prediction model for life expectancy

Open the record for dataset details and reuse information.

publicNov 2022View details →
dryad36/100

Data for: Advances and critical assessment of machine learning techniques for prediction of docking scores

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad36/100

Test data from SPCAM for machine learning in moist physics

Open the record for dataset details and reuse information.

publicMar 2021View details →
dryad36/100

Data from: Training data from SPCAM for machine learning in moist physics

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad36/100

A machine learning approach to integrating genetic and ecological data in tsetse flies (Glossina pallidipes) for spatially explicit vector control planning

Open the record for dataset details and reuse information.

publicOct 2021View details →
dryad36/100

Data for: Foundations of a fast, data-driven, machine-learned simulator

Open the record for dataset details and reuse information.

publicMay 2021View details →
zenodo32/100

Identifying degradation patterns of lithium ion batteries from impedance spectroscopy using machine learning

<p>Dataset accompanying&nbsp;the paper: &quot;Identifying degradation patterns of lithium ion batteries from impedance spectroscopy using machine learning&quot;</p>

opencc-by-4.0Jan 2020View details →
zenodo32/100

Experimental fluvial-deltaic stratigraphic patches for machine learning applications

<p>Collection of 6,132 images&nbsp;(128 x&nbsp;128 pixels) cropped from experimental stratigraphy produced in the Tulane Delta Basin, TDB-10-1, under temporally constant boundary conditions. The images are prepared to be used in a machine learning project.</p> <p>Each image is prefixed with a number [0-5] which indicates the strike section the image was selected from. The cropped strike sections are obtained from the archival dataset located on SEN:&nbsp;<a href="http://sedexp.net/catalog/tdb-10-1-tulane-delta-basin">http://sedexp.net/catalog/tdb-10-1-tulane-delta-basin</a>.</p> <p>After cropping from the strike sections, each image was processed with binarization and a sequence of morphological opening and closing operations. The code that did the processing can be obtained at&nbsp;<a href="https://github.com/amoodie/StratGAN/blob/master/process_images/nrand_process.py">https://github.com/amoodie/StratGAN/blob/master/process_images/nrand_process.py</a>.</p> <p>This data was produced as part of a larger project:&nbsp;<a href="https://github.com/amoodie/StratGAN">https://github.com/amoodie/StratGAN</a></p>

opencc-by-4.0Feb 2020View details →
zenodo32/100

Dataset for "The State of the ML-universe: 10 Years of Artificial Intelligence & Machine Learning Software Development on GitHub"

<p>Supplementary data to &quot;The State of the ML-universe: 10 Years of Artificial Intelligence &amp; Machine Learning Software Development on GitHub&quot; accepted for publication at MSR 2020.</p> <p>The data included in this package were used to conduct analyses to characterize the AI &amp; ML software development community hosted on GitHub. Please read the paper for a full understanding of what data was collected and how it was used.</p> <p>Questions and comments can be directed to Danielle Gonzalez dng2551@rit.edu</p>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Machine learning in mass spectrometry: A MALDI-TOF MS approach to phenotypic antibacterial screening

<p>Dataset relating to the publication:</p> <p>Machine learning in mass spectrometry: A MALDI-TOF MS approach to phenotypic antibacterial screening</p> <p>by Luuk Nico van Oosten and Christian D. Klein</p> <p>Published in the Journal of Medicinal Chemistry, 2020</p> <p><strong>Important notice:</strong></p> <p><strong>The data are free to use for non-commercial, academic purposes, provided that the original source is<br> cited and the authors and the publication are credited in any derivative work.</strong></p> <p><strong>A patent application has been filed for the method described by van Oosten and Klein, which uses mass<br> spectrometry and machine learning to identify the pharmacological or other effects of compounds on cell<br> cultures and other biological systems.</strong></p> <p>Therefore, a license for the commercial use of the method must be negotiated by contacting either</p> <p>Anke Faller<br> Universit&auml;t Heidelberg<br> Dezernat Forschung<br> Rechts- und Strukturfragen der Forschungsf&ouml;rderung<br> Seminarstra&szlig;e 2, 69117 Heidelberg<br> Tel. +49 6221 54-12611<br> anke.faller(at)zuv.uni-heidelberg.de</p> <p>or</p> <p>Prof. Dr. C. Klein; c.klein(at)uni-heidelberg.de<br> Medicinal Chemistry<br> Institute of Pharmacy and Molecular Biotechnology IPMB<br> Heidelberg University, INF 364<br> D-69120 Heidelberg<br> Germany<br> Phone: ++49-6221-54-4875<br> FAX&nbsp; : ++49-6221-54-6430</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Preprocessed Data used by "Machine Learning Enabled Brain Segmentation for Small Animal Image Registration"

<p>Data, preprocessed by the SAMRI package, used to train the models in&nbsp;&quot;Machine Learning Enabled Brain Segmentation for Small Animal Image Registration&quot;. The models can be found <a href="https://zenodo.org/record/3759361#.XrKrgBMzZhE">here</a>.</p>

opencc-by-4.0May 2020View details →
zenodo32/100

Identifying galaxies, quasars and stars with machine learning: a new catalogue of classifications for 111 million SDSS sources without spectra

<p>The Paper:&nbsp;<a href="https://arxiv.org/abs/1909.10963">https://arxiv.org/abs/1909.10963</a></p> <p>Abstract: We used 3.1 million spectroscopically labelled sources from the Sloan Digital Sky Survey (SDSS) to train an optimised random forest classifier using photometry from the SDSS and the Widefield Infrared Survey Explorer (WISE). We applied this machine learning model to 111 million previously unlabelled sources from the SDSS photometric catalogue which did not have existing spectroscopic observations. Our new catalogue contains 50.4 million galaxies, 2.1 million quasars, and 58.8 million stars. We provide individual classification probabilities for each source, with 6.7 million galaxies (13%), 0.33 million quasars (15%), and 41.3 million stars (70%) having classification probabilities greater than 0.99; and 35.1 million galaxies (70%), 0.72 million quasars (34%), and 54.7 million stars (93%) having classification probabilities greater than 0.9. Precision, Recall, and F1 score were determined as a function of selected features and magnitude error. We investigate the effect of class imbalance on our machine learning model and discuss the implications of transfer learning for populations of sources at fainter magnitudes than the training set. We used a non-linear dimension reduction technique (Uniform Manifold Approximation and Projection: UMAP) in unsupervised, semi-supervised, and fully-supervised schemes to visualise the separation of galaxies, quasars, and stars in a two-dimensional space. When applying this algorithm to the 111 million sources without spectra, it is in strong agreement with the class labels applied by our random forest model.</p> <p>When using this dataset, please reference our paper via the journal (<a href="https://arxiv.org/abs/1909.10963">https://arxiv.org/abs/1909.10963</a>) and this DOI (10.5281/zenodo.3459293). If you make use of our scripts please reference our Github repository DOI (10.5281/zenodo.3855160).</p> <p>File descriptions:</p> <p>All of these files are Pandas Dataframes, saved as pickle files. df_spec_classprobs.pkl contains the spectroscopically observed sources used for training and testing. This has been cleaned, and has the results of the random forest classifier added as additional columns (sources used for training have NaNs in the class_pred column). SDSS-ML-all contains the 111 million photometrically observed sources, with our class labels and probabilities added. SDSS-ML-galaxies/quasars/stars is the same file broken up by assigned class for convenience.</p>

opencc-by-4.0Sep 2019View details →
zenodo32/100

Predicting human health from biofluid-based metabolomics using machine learning

<p>Processed input data sets as well as output files and trained health state predictive models for human diseases using metabolomics of biofluids. These data are associated with the following manuscript: &#39;Predicting human health from biofluid-based metabolomics using machine learning&#39; https://doi.org/10.1101/2020.01.29.20019471</p>

opencc-by-4.0Jun 2020View details →
zenodo32/100

Fermilab LHC Physics Center Machine Learning Hands-On Advanced Tutorial Session Datasets

<p>The dataset jet_raw.tar.gz contains a sample of 2011 CMS Open Simulation in numpy arrays. The columns correspond to</p> <pre><code>['run', 'lumi', 'event', 'met', 'sumet', 'rho', 'pthat', 'mcweight', 'njet_ak7', 'jet_pt_ak7', 'jet_eta_ak7', 'jet_phi_ak7', 'jet_E_ak7', 'jet_msd_ak7', 'jet_area_ak7', 'jet_jes_ak7', 'jet_tau21_ak7', 'jet_isW_ak7', 'jet_ncand_ak7', 'ak7pfcand_ijet']</code></pre> <p>Each row is a separate anti-k<sub>T</sub> R=0.7 (AK7)&nbsp;jet. The code to produce the numpy arrays is located at&nbsp;https://doi.org/10.5281/zenodo.3901871</p> <p>The dataset jet_images.h5 contains preprocessed 2D jet images.</p> <p>The datasets ntuple_4mu_bkg.root,&nbsp;ntuple_4mu_gg.root, and&nbsp;ntuple_4mu_VV.root contain simulated LHC events with 4 muons for the background process, gluon fusion Higgs boson production, and vector boson fusion Higgs boson production.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2020View details →
zenodo32/100

Dataset used for detecting DNS over HTTPS by Machine Learning.

<p><strong>&nbsp;</strong>The dataset consists of three different data sources:</p> <ol> <li>&nbsp;DoH enabled Firefox</li> <li>DoH enabled Google Chrome</li> <li>Cloudflared DoH proxy</li> </ol> <p>The capture of web browser data was made using the Selenium framework, which simulated classical user browsing. The browsers received command for visiting domains taken from Alexa&#39;s top 10K most visited websites. The capturing was performed on the host by listening to the network interface of the virtual machine. Overall the dataset contains almost 5,000 web-page visits by Mozilla and 1,000 pages visited by Chrome.</p> <p>The Cloudflared DoH proxy was installed in Raspberry PI, and the IP address of the Raspberry was set as the default DNS resolver in two separate offices in our university. It was continuously capturing the DNS/DoH traffic created up to 20 devices for around three months.</p> <p>The dataset contains 1,128,904 flows from which is around 33,000 labeled as DoH. We provide raw pcap data, CSV with flow data, and CSV file with extracted features.</p> <p>The CSV with extracted features has the following data fields:</p> <p>- Label (1 - Doh, 0 - regular HTTPS)<br> - Data source<br> - Duration<br> - Minimal Inter-Packet Delay<br> - Maximal Inter-Packet Delay<br> - Average Inter-Packet Delay<br> - A variance of Incoming Packet Sizes<br> - A variance of Outgoing Packet Sizes<br> - A ratio of the number of Incoming and outgoing bytes<br> - A ration of the number of Incoming and outgoing packets<br> - Average of Incoming Packet sizes<br> - Average of Outgoing Packet sizes<br> - The median value of Incoming Packet sizes<br> - The median value of outgoing Packet sizes<br> - The ratio of bursts and pauses<br> - Number of bursts<br> - Number of pauses<br> - Autocorrelation<br> - Transmission symmetry in the 1st third of connection<br> - Transmission symmetry in the 2nd third of connection<br> - Transmission symmetry in the last third of connection</p> <p>The observed network traffic does not contain privacy-sensitive information.&nbsp;</p> <p>The zip file structure is:</p> <pre><code>|-- data |   |-- extracted-features...extracted features used in ML for DoH recognition |   |   |-- chrome |   |   |-- cloudflared |   |   `-- firefox |   |-- flows...............................................exported flow data |   |   |-- chrome |   |   |-- cloudflared |   |   `-- firefox |   `-- pcaps....................................................raw PCAP data |       |-- chrome |       |-- cloudflared |       `-- firefox |-- LICENSE `-- README.md</code></pre> <p><br> When using this dataset, please cite the original work as follows:</p> <pre><code>@inproceedings{vekshin2020, author = {Vekshin, Dmitrii and Hynek, Karel and Cejka, Tomas}, title = {DoH Insight: Detecting DNS over HTTPS by Machine Learning}, year = {2020}, isbn = {9781450388337}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3407023.3409192}, doi = {10.1145/3407023.3409192}, booktitle = {Proceedings of the 15th International Conference on Availability, Reliability and Security}, articleno = {87}, numpages = {8}, keywords = {classification, DoH, DNS over HTTPS, machine learning, detection, datasets}, location = {Virtual Event, Ireland}, series = {ARES '20} } </code></pre> <p>&nbsp;</p>

openmit-licenseMay 2020View details →
dryad32/100

Data from: ClinicNet: machine learning for personalized order set recommendations

<p class="MsoTitle"><b>Objective</b></p> <p>This study assesses whether neural networks trained on electronic health record (EHR) data can anticipate what individual clinical orders and existing institutional order set templates clinicians will use more accurately than existing decision support tools.</p> <p><b>Materials and Methods</b></p> <p>We process 57,624 patients-worth of clinical event EHR data from 2008-2014. We train a feed-forward neural network (ClinicNet) and logistic regression applied to <span>the traditional problem structure of predicting individual clinical items  as well as our proposed workflow of predicting existing institutional order set template usage. </span></p> <p><b>Results</b></p> <p>ClinicNet predicts individual clinical orders (precision=0.32, recall=0.47) better than existing institutional order sets (precision=0.15, recall=0.46). The ClinicNet model predicts clinician usage of existing institutional order sets (avg. precision=0.31) with higher average precision than  a baseline of order set usage frequencies (avg. precision=0.20) or a logistic regression model (avg. precision=0.12).</p> <p><b>Discussion</b></p> <p>Machine learning methods can predict clinical decision-making patterns with greater accuracy and less manual effort than existing static order set templates. This can streamline existing clinical workflows, but may not fit if historical clinical ordering practices are incorrect. For this reason, manually authored content such as order set templates remain valuable for purposeful design of care pathways. ClinicNet's capability of predicting such personalized order set templates illustrates the potential of combining both top-down and bottom-up approaches to delivering clinical decision support content.</p> <p><b>Conclusion</b></p> <p>ClinicNet illustrates the capability for machine learning methods applied to the EHR to anticipate both individual clinical orders and existing order set templates, which has the potential to improve upon current standards of practice in clinical order entry.</p>

opencc-zeroJun 2020View details →
dryad32/100

Data from: QTG-Finder2: a generalized machine-learning algorithm for prioritizing QTL causal genes in plants

Linkage mapping has been widely used to identify quantitative trait loci (QTL) in many plants and usually requires a time-consuming and labor-intensive fine mapping process to find the causal gene underlying the QTL. Previously, we described QTG-Finder, a machine-learning algorithm to rationally prioritize candidate causal genes in QTLs. While it showed good performance, QTG-Finder could only be used in Arabidopsis and rice because of the limited number of known causal genes in other species. Here we tested the feasibility of enabling QTG-Finder to work on species that have few or no known causal genes by using orthologs of known causal genes as training set. The model trained with orthologs could recall about 64% of Arabidopsis and 83% of rice causal genes when the top 20% ranked genes were considered, which is similar to the performance of models trained with known causal genes. The average precision was 0.027 for Arabidopsis and 0.029 for rice. We further extended the algorithm to include polymorphisms in conserved non-coding sequences and gene presence/absence variation as additional features. Using this algorithm, QTG-Finder2, we trained and cross-validated Sorghum bicolor and Setaria viridis models. The S. bicolor model was validated by causal genes curated from the literature and could recall 70% of causal genes when the top 20% ranked genes were considered. In addition, we applied the S. viridis model and public transcriptome data to prioritize a plant height QTL and identified 13 candidate genes. QTL-Finder2 can accelerate the discovery of causal genes in any plant species and facilitate agricultural trait improvement.

opencc-zeroAug 2020View details →
zenodo32/100

Data and scripts for 'Potential and limitations of machine learning for modeling warm-rain cloud microphysical processes'

<p>Data and scripts&nbsp;for &quot;Potential and limitations of machine learning for modeling warm-rain cloud microphysical processes&quot; by Axel Seifert and Stephan Rasp,&nbsp;J. Adv. Modeling Earth Systems, 12, 2020, https://doi.org/10.1029/2020MS002301</p>

opencc-by-4.0Aug 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record