Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
251
datasets available to search
ShareScore release 0.7.1
Dataset results
251 results for “deep learning models”
Replication Package for "PyTraceBERT: Python Traceback-based Language Model for Detecting Compatibility Issues in Deep Learning Systems"
<p>This package contains the traceback data, pre-trained models, and static word embeddings used in the paper, PyTraceBERT: Python Traceback-based Language Model for Detecting Compatibility Issues in Deep Learning Systems.</p>
Winter Precipitation-Type Models for "Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications"
<p>This contains trained model weights, scalers, and evaluation metrics for the winter precipitation-type models trained as part of the paper "Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications". </p>
Enhancing Smartphone Battery Life: A Deep Learning Model Based on User-Specific Application and Network Behaviour
<p>This work presents an analysis based on training AI models directly on devices to make personalized predictions tailored to individual usage patterns, ensuring that each user benefits from a personalized approach to battery management. By integrating these AI-based insights, mobile devices can proactively manage power consumption, improving battery performance and user satisfaction. This personalized, intelligent approach to battery management represents a significant advance in optimizing device efficiency and addresses the growing demand for longer-lasting mobile technology.</p>
Deep learning model for characterizing protein-RNA interactions from sequence at single-base resolution
<p> </p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip.h5</a> - This file contains the training, validation, and test data for the Reformer model.</p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip_bc.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip_bc.h5</a> - This file contains the training, validation, and test data for the Reformer-BC model.</p> <p><a href="https://zenodo.org/api/records/14027315/draft/files/Reformer-code.zip/content" target="_blank" rel="noopener">Reformer-code.zip</a> - This file contains the training code of Reformer.</p>
Data for "Deep learning-based model for diagnosing Alzheimer's disease and tauopathies"
<p>Image datasets and tuned models used in the paper (Koga et al., 2021). Data.zip contains image and text files for training models. Test.zip contains 12 images from 4 patients, which are a part of the hold-out dataset images used in the paper. There are 9 CSV files, which contain the results of tau burden quantification. Python code is available at GitHub (<a href="https://github.com/Koga-MD/DL-Tauopathies">https://github.com/Koga-MD/DL-Tauopathies</a>). </p>
CNN models and training, validation and test datasets for "PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data"
<p>Convolutional neural network (CNN) models and their respective training, validation and test datasets used in manuscript:</p> <p>Tuomo Hartonen, Teemu Kivioja and Jussi Taipale, "PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data"</p>
Training Deep Learning Models to Estimate SWAT Parameters using Streamflow Observations
<p>This folder provides the simulation and observational data</p> <p>Simulation data using SWAT (1000 realz)<br> Train, Val, and Test splits (80/10/10)<br> Observational data for ARW (WY2000-2016)</p>
Datasets used to train the models in "Deep learning for denoising High-Rate Global Navigation Satellite System data."
<p>Datasets used to train the models in "Deep learning for denoising High-Rate Global Navigation Satellite System data." Additional information can be found at https://github.com/amtseismo/hrgnss_denoising.</p>
Deep learning models challenge the prevailing assumption that face-like effects for objects of expertise support domain-general mechanisms
<p>The question of whether perceptual expertise is mediated by general-expert or domain-specific processing mechanisms has been debated for decades. Because humans are experts in face recognition, face-like neural and cognitive effects for objects of expertise were considered to support for the general-expertise hypothesis. Conversely, stronger effects for faces than objects of expertise were considered to support the domain-specific hypothesis. However, the effects of domain, experience, and level of categorization, are confounded in human studies, which may lead to erroneous inferences. To overcome these limitations, we used computational models of perceptual expertise and tested different domains (objects, faces, birds) and levels of categorization (basic, sub-ordinate, individual) in isolation, matched for amount of experience. Like humans, the models generated a larger inversion effect for faces than for objects. Importantly, a face-like inversion effect was found for individual-based categorization of non-faces (birds) but only in a network specialized for that domain. Thus, contrary to prevalent assumptions, face-like effects in objects of expertise may originate from domain-specific rather than domain-general processing mechanisms. More generally, we show how deep learning algorithms can be used to isolate the effects of factors that are inherently confounded in the natural environment of biological organisms.</p>
Training patches and prediction codes of deep learning (LANA) model for Landsat 8/9 cloud/shadow mask
<p>This dataset includes (i) the image patches dataset and (ii) application/prediction (not training) codes for Landsat 8 cloud and cloud shadow masking used in a paper in review and uploaded here: </p> <p>Hankui Zhang, Dong Luo, David Roy, A learning attention network algorithm (LANA) for accurate Landsat-8 cloud and shadow masking, <em>Remote Sensing of Environment</em> </p> <p>The documentation is in <a href="https://zenodo.org/api/files/5462baa5-2bba-4b0f-92aa-c17681b6464b/l8_training_data_readme_new.pdf?versionId=95253efb-6447-46d8-ae4b-ac0c93b43532">l8_training_data_readme_new.pdf</a>. </p>
Image dataset for cow identification, including code to train deep learning model, as well as analysis of results (SmARtview, 51088)
<p>This dataset and code was a result of the UKRI project "SmARtview: An AI-powered Augmented Reality Tool for Animal Health and Productivity", linked here: <a href="https://gtr.ukri.org/projects?ref=51088">https://gtr.ukri.org/projects?ref=51088</a></p> <p>These files are intended to be used for an accompanying publication in an academic journal.</p> <p>Anyone is free to use the contents for research and teaching purposes.</p>
Deep learning models challenge the prevailing assumption that face-like effects for objects of expertise support domain-general mechanisms
Open the record for dataset details and reuse information.
phyddle: Software for exploring phylogenetic models with deep learning
Open the record for dataset details and reuse information.
Data and code from: Learning a deep language model for microbiomes: The power of large scale unlabeled microbiome data
Open the record for dataset details and reuse information.
Benign samples used in article "DeepDetectNet vs RLAttackNet: An Adversarial Method to Improve Deep Learning-based Static Malware Detection Model"
<p>This repository contains all benign samples used in article "DeepDetectNet vs RLAttackNet: An Adversarial Method to Improve Deep Learning-based Static Malware Detection Model". It is safe to download these samples.</p>
Anonymized Dataset for "Towards a Better Understanding of Reverse-Complement Equivariance for Deep Learning Models in Genomics"
<p>Anonymous dataset for the paper "Towards a Better Understanding of Reverse-Complement Equivariance for Deep Learning Models in Genomics." Includes data for simulated, binary prediction, and profile prediction tasks. </p>
Data and trained word2vec model for ``Easy over Hard: A Case Study on Deep Learning''
<p>The data include: training and testing data pairs</p> <p>The word2vec model is pre-trained. </p> <p>More details, please refer to the paper</p>
Pre-trained word2vec models for ``Easy over Hard: A Case Study on Deep Learning''
<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds.</p> <p> </p> <p>More details about how to use it, please see paper </p>
SynProtX: A Large-Scale Proteomics-Based Deep Learning Model for Predicting Synergistic Anticancer Drug Combinations
<h2>SynProtX: A Large-Scale Proteomics-Based Deep Learning Model for Predicting Synergistic Anticancer Drug Combinations</h2> <p>SynProtX is a deep learning model that integrates large-scale proteomics data, molecular graphs, and chemical fingerprints to predict synergistic effects of anticancer drug combinations. It provides robust performance across tissue-specific and study-specific datasets, enhancing reproducibility and biological relevance in drug synergy prediction.</p> <p>This Zenodo repository includes a <code>.tar.gz</code> archive containing all essential components to reproduce the experiments described in the study. This archive is designed to work seamlessly with the coding pipeline available at: <a href="https://github.com/manbaritone/SynProtX" target="_blank" rel="noopener">https://github.com/manbaritone/SynProtX</a>.</p> <h3>License:</h3> <p>Creative Commons Zero v1.0 Universal (CC0)<br>This work is released under CC0, dedicating it to the public domain. You are free to use, modify, and distribute it without restriction.</p> <h3>Archive Contents:</h3> <p>This compressed file includes:</p> <ul> <li>Datasets<br>- Tissue Datasets: <code>ALMANAC-Breast</code>, <code>ALMANAC-Lung</code>, <code>ALMANAC-Ovary</code>, <code>ALMANAC-Skin</code><br>- Study Datasets: <code>FRIEDMAN</code>, <code>ONEIL</code></li> <li>Supporting Files<br>- Raw and preprocessed data<br>- Feature dictionaries<br>- Hyperparameter configurations<br>- Trained model weights</li> </ul> <h3>Folder Structure:</h3> <blockquote> <p><code>SynProtX/</code><br><code>├── data/ # Raw and preprocessed data</code><br><code>│ ├── export/ # Processed protein/gene expression & drug combinations</code><br><code>│ ├── nps/ # Numpy arrays for all datasets</code><br><code>│ ├── nps_intersected/ # Dataset-specific numpy arrays</code><br><code>│ └── raw/ # Original data from DrugComb, CCLE, COSMIC, ChEMBL V31, ProCan-DepMapSanger</code><br><code>├── feature_dicts/ # Feature dictionaries for drug combinations</code><br><code>├── hyperparams/ # Hyperparameter configs for SynProtX-GATFP</code><br><code>│ ├── classification/ # For classification tasks</code><br><code>│ └── regression/ # For regression tasks</code><br><code>├── state_dict/ # Trained model weights</code><br><code>│ ├── classification/ # PyTorch checkpoints for classification</code><br><code>│ └── regression/ # PyTorch checkpoints for regression</code><br><code>└── README_Zenodo.md # This file</code></p> </blockquote> <h3>For more information, please visit:</h3> <p><strong>GitHub:</strong> <a href="https://github.com/manbaritone/SynProtX" target="_blank" rel="noopener">https://github.com/manbaritone/SynProtX</a></p>
DeepAnnotation: A novel interpretable deep learning-based genomic selection model that integrates comprehensive functional annotations
<p>1. Update package, example dataset, and demo code of DeepAnnotation</p> <p>2. Update the transformed genotype data, the phenotype data, the comprehensive functional annotation data for Duroc prepared by RNAfold, DeepSEA, easyMF models, and the four types of input data for training DeepAnnotation model</p> <p>3. Add the conserved functional annotation</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.