Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

32

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

32 results for “deep learning; machine learning”

Learn how ShareScore rates datasets ↗
zenodo40/100

Companion for "Understanding Distributed Deep Learning Performance by Correlating HPC and Machine Learning Measurements"

<p>This is the Companion Material for the paper &ldquo;Understanding Distributed Deep Learning Performance by Correlating HPC and Machine Learning Measurements&rdquo;, by Ana Luisa Veroneze Sol&oacute;rzano and Lucas Mello Schnorr. The manuscript was approved for publication in the <a href="https://www.isc-hpc.com/research-papers-2022.html">ISC High Performance 2022</a>&nbsp;for the Research Papers session. A public companion is also availabl in GitLab:&nbsp;<a href="https://gitlab.com/anaveroneze/isc2022-companion/">https://gitlab.com/anaveroneze/isc2022-companion</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Cloud to Thing Continuum based Sports Monitoring System using Machine Learning and Deep Learning Model

<p><span>Sports monitoring and analysis have seen significant advancements with the integration of cloud computing and continuum paradigms, facilitated by machine learning and deep learning techniques. In this study, we present a novel approach for sports monitoring that seamlessly transitions from traditional cloud-based architectures to a continuum paradigm, enabling real-time analysis and insights into player performance and team dynamics. Leveraging machine learning and deep learning algorithms, our framework offers enhanced capabilities for player tracking, action recognition, and performance evaluation in various sports scenarios. This research proposes a Cloud-to-Thing Continuum based Sports Monitoring System utilizing Machine Learning (ML) and Deep Learning (DL) models. The system integrates data acquisition, preprocessing, feature extraction, cloud-based processing, continuum paradigm integration, and decision-making stages. It leverages innovative techniques such as Improved Mask R-CNN for pose estimation, hybrid metaheuristic algorithms with Generative Adversarial Network (GAN) for classification, and fuzzy decision-making Based on the integrated analysis, decisions are made regarding player performance, team strategies, and tactical adjustments. The continuum approach ensures a balance between centralized cloud processing and distributed edge processing, optimizing resource utilization and reducing latency. Through this system, real-time analysis of sports events is achieved, enabling immediate feedback for time-sensitive applications.</span></p>

opencc-by-4.0May 2024View details →
zenodo40/100

Data related to the publication "Efficient molecular dynamics simulations of deep eutectic solvents with first-principles accuracy using machine learning interatomic potentials"

<p>The training data sets, the trained machine learning models, and input scripts for the training and molecular dynamics simulations.</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

An Artificial Eye for Palaeography. Applying Deep Machine Learning for the Study of Medieval Latin Scripts

<p>The project &ldquo;Digital Forensics for Historical Documents&rdquo; (at Huygens ING, Amsterdam) attempts to create a digital tool, based on a deep learning system, in which the unique characteristics of one medieval script sample will be matched with similar script samples by making use of digitized manuscript collections available in the world wide web.</p> <p>Project website and contact: <a href="https://www.youtube.com/redirect?q=https%3A%2F%2Fen.huygens.knaw.nl%2Fprojecten%2Fdigital-forensics-for-historical-documents%2F&amp;v=WYtseNK-1Dc&amp;event=video_description&amp;redir_token=QUFFLUhqbDM0WDRJRjN0V3QzOXV6d0ZNbDB2TVYzV1hUQXxBQ3Jtc0ttLVJHMEE1RlRkZjJVV3poYnpQOHZseTFzZHVqWHRGR2k2eWpVWTJSSldtV2p4aFhKbFRIQTFjVHQ5YnM0Mkd4ajBzOTNtdE1yOVJLVktqWlhqWUgzZXc3YmNQMS1nNGFpZ3p2amNOTFNqTDVUSTdtZw%3D%3D">https://en.huygens.knaw.nl/projecten/...</a>&nbsp;</p> <p>Presented as a Lightning Talk for the Schoenberg Symposium 2020</p>

opencc-by-4.0Nov 2020View details →
dryad36/100

The Camouflage Machine: Optimising protective colouration using deep learning with genetic algorithms

Evolutionary biologists frequently wish to measure the fitness of alternative phenotypes using behavioural experiments. However, many phenotypes are complex. For example colouration: camouflage aims to make detection harder, while conspicuous signals (e.g. for warning or mate attraction) require the opposite. Identifying the hardest and easiest to find patterns is essential for understanding the evolutionary forces that shape protective colouration, but the parameter space of potential patterns (coloured visual textures) is vast, limiting previous empirical studies to a narrow range of phenotypes. Here we demonstrate how deep learning combined with genetic algorithms can be used to augment behavioural experiments, identifying both the best camouflage and the most conspicuous signal(s) from an arbitrarily vast array of patterns. To show the generality of our approach, we do so for both trichromatic (e.g. human) and dichromat (e.g. typical mammalian) visual systems, in two different habitats. The patterns identified were validated using human participants; those identified as the best for camouflage were significantly harder to find than a tried-and-tested military design, while those identified as most conspicuous were significantly easier than other patterns. More generally, our method, dubbed the 'Camouflage Machine', will be a useful tool for identifying the optimal phenotype in high dimensional state-spaces.

opencc-zeroDec 2020View details →
zenodo36/100

Raw NGS Data for "Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins"

<p>This directory contains relevant fastq files used for deep sequencing analysis in the publication &ldquo;Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins&rdquo;.&nbsp;</p> <p>Fastq files are provided for presorted, uninduced and induced populations from DMS experiments of&nbsp;four&nbsp;homologs (TtgR, TetR, RolR, and MphR). Three replicates were performed for each sample.</p> <p>Data analysis of this&nbsp;deep sequencing data was performed using custom scripts, which are described in the methods section of the publication.</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Artificial Intelligence and COVID-19 using chest CT scan and chest X-ray images: Machine Learning and Deep Learning Approaches for Diagnosis and Treatment

<p>We uploaded the Table of included articles in the systematic&nbsp; review &quot;Artificial Intelligence and COVID-19 using chest CT scan and chest X-ray images: Machine Learning and Deep Learning Approaches for Diagnosis and Treatment&quot;</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Ensemble BLUP, Machine Learning, and Deep Learning Models Predict Maize Yield Better Than Each Model Alone.

<p>Data and scripts exploring ensembling strategies using the models developed in <a href="https://academic.oup.com/g3journal/advance-article/doi/10.1093/g3journal/jkad006/6982634">Kick et al., 2023</a> (see also <a href="https://zenodo.org/record/7401113">1</a>, <a href="https://zenodo.org/record/6916775">2</a>). Download all files to a single directory then run setup.sh or manually unzip using tar.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>Filename</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>setup.sh</td> <td>Simple script that unzips zipped directories</td> </tr> <tr> <td>ext_data</td> <td>Reduced data from Kick et al. 2023</td> </tr> <tr> <td>ext_data_notebooks</td> <td>Contains python notebooks containing analysis and R markdown file containing visualization of results. Python and R data objects are written to allow results to be read in instead of re-generated.</td> </tr> <tr> <td>output</td> <td>Folder containing a placeholder file.</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>This research used resources provided by the United States Department of Agriculture&rsquo;s Agricultural Research Service (project number 5070-21000-041-000-D). The SCINet project of the USDA Agricultural Research Service (project number 0500-00093-001-00-D) was instrumental in the training of the models used in this work. In addition, we would like to acknowledge those presently and historically involved in generating data for the Genomes to Fields Initiative.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-3.0-usMar 2023View details →
dryad36/100

The Camouflage Machine: Optimising protective colouration using deep learning with genetic algorithms

Open the record for dataset details and reuse information.

publicDec 2020View details →
zenodo32/100

Blinded Predictions and Post-hoc Analysis of the Second Solubility Challenge Data: Exploring Training Data and Feature Set Selection for Machine and Deep Learning Models

<p>Training and test datasets and scripts for training models.</p>

openmit-licenseSep 2022View details →
zenodo32/100

Dataset and results for "Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation"

<p>Dataset and results for &quot;Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation&quot;</p> <p>Yuhang Zhang1, Aizhong Ye1*, Phu Nguyen2, Bita Analui2, Soroosh Sorooshian2, Kuolin Hsu2</p> <p>1 State Key Laboratory of Earth Surface Processes and Resource Ecology, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China.</p> <p>2 Center for Hydrometeorology and Remote Sensing, Department of Civil and Environmental Engineering, University of California, Irvine, Irvine, California, CA 92697, USA.</p> <p>## Dataset&nbsp;&nbsp; &nbsp;</p> <p>Streamflow simulations from one observed precipitation (CMA) and three satellite precipitation products (PDIR, IMERG-F, and GSMaP) for 522 sub-basins.</p> <p>- Q-CMA (streamflow reference)<br> - Q-PDIR (uncorrected)<br> - Q-IMERGF (uncorrected)<br> - Q-GSMAP (uncorrected)</p> <p>### Data structure</p> <p>- Head section (row1-row5)<br> &nbsp; - SubNO:&nbsp;&nbsp; &nbsp;522&nbsp;<br> &nbsp; - BeginT:&nbsp;&nbsp; &nbsp;2003-01-01 00:00&nbsp;<br> &nbsp; - EndT:&nbsp;&nbsp; &nbsp;2019-12-31 00:00&nbsp;<br> &nbsp; - Interval:&nbsp;&nbsp; &nbsp;1440s (daily)<br> &nbsp; - Revise:&nbsp;&nbsp; &nbsp;10 (scaling factor to keep int datatype)<br> &nbsp; - Point1&nbsp;&nbsp; &nbsp;Point2&nbsp;&nbsp; &nbsp;... (Subbasin No.)<br> - Data section<br> &nbsp; - 6209 rows, 522 cols</p> <p>## Results</p> <p>Two post-processing model results for test period (2015-1-1 to 2018-12-31).</p> <p>### Data structure</p> <p>- 1462 rows, every row denotes each day from 2015-1-1 to 2018-12-31</p> <p>- 100 columns, every column denotes each quantile from 0.005 to 0.995, total 100 quantiles.</p> <p>### qrf-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p>### lstm-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo32/100

Data related to the publication "structure and transport properties of LiTFSI-based deep eutectic electrolytes from machine-learned interatomic potential simulations"

<p>Reference training and test datasets, trained ML potential models, and input scripts for the training (Allegro) and MD simulations (LAMMPS).</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Dataset for manuscript "Turnover number predictions for kinetically uncharacterized enzymes using machine and deep learning"

<p>Dataset for the github repository containing the code for the manuscript &quot;Turnover number predictions for kinetically uncharacterized enzymes using machine and deep learning&quot;.</p>

opencc-by-4.0Nov 2022View details →
ClinicalTrials.gov32/100

Deciphering AMD by Deep Phenotyping and Machine Learning- Prospective Study - PINNACLE

ClinicalTrials.gov study NCT04269304. IPD Sharing: YES. Countries: 3. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov32/100

Benefit of Machine Learning to Diagnose Deep Vein Thrombosis Compared to Gold Standard Ultrasound

ClinicalTrials.gov study NCT05288413. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

EEG Recordings and Analysis in Parkinson's Patients: Towards Adaptive Deep Brain Stimulation by Machine Learning

ClinicalTrials.gov study NCT05284526. IPD Sharing: Not stated. Countries: 1. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

Machine Learning on the Impacts of Mutations in the SARS-CoV-2 Spike RBD on Binding Affinity to Human ACE2 based on Deep Mutational Scanning Data

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
ClinicalTrials.gov28/100

Machine and Deep Learning for Congenital Diaphragmatic Hernia (CLANNISH)

ClinicalTrials.gov study NCT04609163. IPD Sharing: Not stated. Countries: 0. Publications: 3.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov28/100

Fall Risk Assessment Using Hybrid Machine Learning and Deep Learning Approaches and a Novel Posturography

ClinicalTrials.gov study NCT05308563. IPD Sharing: Not stated. Countries: 0. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
geo24/100

Deep and accurate detection of m6A RNA modifications in human and mouse cells using miCLIP2 and m6Aboost machine learning [b]

GEO Series GSE163492. Homo sapiens. 4 samples. Type: Other.

openGEO-OpenJun 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record