Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

251

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

251 results for “deep learning models”

Learn how ShareScore rates datasets ↗
zenodo40/100

A Deep Learning-Based Hybrid Model of Global Terrestrial Evaporation

<p>This repository contains the datasets used in the research article &quot;A Deep Learning-Based Hybrid Model of Global Terrestrial Evaporation&quot;.</p> <p>The repository contains the following files: 1) Input - contains all the processed input used for training the deep learning models and the datasets used for creating the figures in the article. 2) Output - contains the final deep learning models and the outputs (evaporation and transpiration stress factor) outputs from the hybrid model developed in the study.</p> <p>Formats: All scripts are in the programming language Python. The datasets are in HDF5 and NetCDF file formats.</p> <p>The codes related to the research article and deep learning model are available in the following repository: https://github.com/akashkoppa/StressNet</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Supplementary Dataset for Deep learning based kcat prediction enables improved enzyme constrained model reconstruction

<p>This dataset is the supplementary dataset for the paper &quot;<strong>Deep learning based&nbsp;<em>k</em><sub>cat</sub>&nbsp;prediction enables improved enzyme constrained model reconstruction</strong>&quot;. Protein sequence fasta files, deep learning predicted&nbsp;<em>k</em><sub>cat</sub>&nbsp;values, classcial-ecGEMs, DL-ecGEMs and&nbsp;<em>Posterior</em>-mean-ecGEMs for 343 yeast/fungi species are available in this dataset.This repository also contains the computed results for reproducing the figures as model_build_files .&nbsp;The scripts can be found in&nbsp;Github (https://github.com/SysBioChalmers/DLKcat)</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Deep Learning for Reaction-Diffusion Glioma Growth Modeling: Towards a Fully Personalized Model? — Supporting Data

<p>Supporting data for&nbsp;Martens et al. Deep Learning for Reaction-Diffusion Glioma Growth Modelling: Towards a Fully Personalised Model? arXiv:2111.13404.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Dataset of very-high-resolution satellite RGB images to train deep learning models to detect and segment high-mountain juniper shrubs in Sierra Nevada (Spain)

<p>This dataset provides annotated very-high-resolution satellite RGB images extracted from Google Earth to train deep learning models to perform instance segmentation of Juniperus communis L. and Juniperus sabina L. shrubs. All images are from the high mountain of Sierra Nevada in Spain. The dataset contains 810 images (.jpg) of size 224x224 pixels. We also provide partitioning of the data into Train (567 images), Test (162 images), and Validation (81 images) subsets. Their annotations are provided in three different .json files following the COCO annotation format.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Dataset of very-high-resolution satellite RGB images to train deep learning models to recognize high-mountain juniper shrubs from Sierra Nevada (Spain)

<p>This dataset provides annotated very-high-resolution satellite RGB images extracted from Google Earth to train deep learning models to recognize Juniperus communis L. and Juniperus sabina L. shrubs.&nbsp; All images are from the high mountain of Sierra Nevada in Spain. The dataset contains 2000 images (.jpg) of size 512x512 pixels partitioned into two classes: Shrubs and NoShrubs. We also provide partitioning of the data into Train (1800 images), Test (100 images), and Validation (100 images) subsets.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

A cooperative deep learning model for stock market prediction using deep autoencoder and sentiment analysis

<p>This data is used for Stock Market Prediction.&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Deep Learning based Urban Morphology for City-scale Environmental Modeling

<p>The WRF simulations were performed using the Weather Research and Forecasting (WRF) model, version 4.2.1. The three nested domains are centered over Chicago, USA, with a spatial resolution of 9, 3, and 1 km for the outermost, middle, and innermost domains. The model was implemented with 42 pressure levels, with the first model level located at 21.2 m and the first 1 km vertical height containing 11 model levels. The initial and boundary conditions are taken from the National Centers for Environmental Prediction (NCEP) Final Reanalysis dataset at 1 degree spatial and 6-hourly temporal resolution.</p><p>The physics components include the WRF single moment 6 class for microphysics, Dudhia for shortwave, the Rapid Radiative Transfer Model for longwave radiation parameterizations, Bougeault for the planetary boundary layer, Noah for the land surface model, Building Environment Parametrization (BEP)&nbsp; for the urban model, and Grell for the cumulus scheme (only for the outermost domain of 9 km spatial resolution). The LCZs of Chicago, USA, are generated using the crowd-sourcing method. The training dataset, created manually, is obtained from the WUDAPT portal, and random forest classification is applied to Landsat 8 imagery to derive the LCZs for the desired region. The simulations are performed from 1/Jul/2018 00:00 to 7/Jul/2018 06:00, where the first 6 hours are discarded as spin-up time.</p><p>The Digital Synthetic City (DSC) of Chicago, USA, uses satellite imagery and global-scale population and elevation data as input to the automatic method for producing a statistically similar and synthetic city-scale 3D urban model as output.</p><p>The Control simulations use National Land Cover Database land use/land cover with&nbsp; NUDAPT parameters, the three default WRF urban classes, and corresponding UCPs; the WUDAPT uses the MODIS classes with additional urban LCZs and UCPs from Brousse et al. (2016), and the DSC uses the WUDAPT classes with UCPs generated from DSC method.</p><p>The dataset contains:</p><p>1. Output from DSC in Shapefile.</p><p>2. WRF model output for the third domain (1 km) spatial resolution domain for (a) NUDAPT or Control (b) WUDAPT or LCZs (c) DSC</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

SCM-CNN: A Robust Deep Learning Matting Model for Cloud Removal in Optical Imagery

<p>This is a dataset that can be used for cloud detection and cloud opacity estimation. The data is saved in python-numpy form, and stored in dictionary :dict_keys([&#39;OriginImage&#39;, &#39;Gimage&#39;, &#39;Alpha&#39;, &#39;Trimap&#39;, &#39;CloudMaxDN&#39;]) represents the cloud-free remote sensing image, cloud remote sensing image, cloud opacity, trilateration information and cloud brightness respectively. The command {np.load(&quot;Path&quot;,allow_pickle=True).item()} is used to read, where &quot;Path&quot; is the corresponding path to the file.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Fig. 9 in An optimised YOLOv4 deep learning model for efficient malarial cell detection in thin blood smear images

Fig. 9 Overall framework of proposed automated malaria diagnosis and species identification. CNN, Convolutional neural network; RBC, red blood cell; YOLO, You Only Look Once (model)

opencc-by-4.0Apr 2024View details →
zenodo40/100

Fig. 8 in An optimised YOLOv4 deep learning model for efficient malarial cell detection in thin blood smear images

Fig. 8 Examples of false positive predictions by the YOLOv4-RC3_4 model. YOLO, You Only Look Once (model)

opencc-by-4.0Apr 2024View details →
zenodo40/100

Fig. 6 in An optimised YOLOv4 deep learning model for efficient malarial cell detection in thin blood smear images

Fig. 6 Comparison of detection performance by the original YOLOv4 model and the YOLOv4-RC3_4 model. Red arrows indicate cells not detected by the original YOLOv4 model, green arrows indicate the same cells detected by the YOLOv4-RC3_4 model. YOLO, You Only Look Once (model)

opencc-by-4.0Apr 2024View details →
zenodo40/100

Fig. 3 in An optimised YOLOv4 deep learning model for efficient malarial cell detection in thin blood smear images

Fig. 3 Network structure of YOLOv4. CSP, cross-spatial connection; SPP, spatial pyramid pooling layer; PANet, Path Aggregation Network; CBM, Convolutional, Batch Normalisation, and Activation; CBL, Convolutional, Batch normalisation, and Leaky-ReLU; Conv, convolutional; Concat, concatenation

opencc-by-4.0Apr 2024View details →
zenodo40/100

Fig. 4 in An optimised YOLOv4 deep learning model for efficient malarial cell detection in thin blood smear images

Fig. 4 Building blocks of the residual learning module. CBM, Convolutional, Batch normalisation and Mish (modules)

opencc-by-4.0Apr 2024View details →
zenodo40/100

Fig. 5 in An optimised YOLOv4 deep learning model for efficient malarial cell detection in thin blood smear images

Fig. 5 Visual representation of the removal of residual blocks from C3 and C4 Res-block body. YOLO,You Only Look Once (model)

opencc-by-4.0Apr 2024View details →
zenodo40/100

Fig. 1 in An optimised YOLOv4 deep learning model for efficient malarial cell detection in thin blood smear images

Fig. 1 Comparison of malaria diagnosis using deep learning CNN models and deep learning object detectors. CNN, Convolutional neural network

opencc-by-4.0Apr 2024View details →
zenodo40/100

Fig. 2 in An optimised YOLOv4 deep learning model for efficient malarial cell detection in thin blood smear images

Fig. 2 Cropping of infected cells using the coordinates of predictions by the object detectors. RBC, Red blood cell; YOLO,You Only Look Once (model)

opencc-by-4.0Apr 2024View details →
zenodo40/100

Cloud to Thing Continuum based Sports Monitoring System using Machine Learning and Deep Learning Model

<p><span>Sports monitoring and analysis have seen significant advancements with the integration of cloud computing and continuum paradigms, facilitated by machine learning and deep learning techniques. In this study, we present a novel approach for sports monitoring that seamlessly transitions from traditional cloud-based architectures to a continuum paradigm, enabling real-time analysis and insights into player performance and team dynamics. Leveraging machine learning and deep learning algorithms, our framework offers enhanced capabilities for player tracking, action recognition, and performance evaluation in various sports scenarios. This research proposes a Cloud-to-Thing Continuum based Sports Monitoring System utilizing Machine Learning (ML) and Deep Learning (DL) models. The system integrates data acquisition, preprocessing, feature extraction, cloud-based processing, continuum paradigm integration, and decision-making stages. It leverages innovative techniques such as Improved Mask R-CNN for pose estimation, hybrid metaheuristic algorithms with Generative Adversarial Network (GAN) for classification, and fuzzy decision-making Based on the integrated analysis, decisions are made regarding player performance, team strategies, and tactical adjustments. The continuum approach ensures a balance between centralized cloud processing and distributed edge processing, optimizing resource utilization and reducing latency. Through this system, real-time analysis of sports events is achieved, enabling immediate feedback for time-sensitive applications.</span></p>

opencc-by-4.0May 2024View details →
dryad40/100

Data and code from: Learning a deep language model for microbiomes: The power of large scale unlabeled microbiome data

<p>We use open source human gut microbiome data to learn a microbial "language" model by adapting techniques from Natural Language Processing (NLP). Our microbial "language" model is trained in a self-supervised fashion (i.e., without additional external labels) to capture the interactions among different microbial species and the common compositional patterns in microbial communities. The learned model produces contextualized taxa representations that allow a single bacteria species to be represented differently according to the specific microbial environment it appears in. The model further provides a sample representation by collectively interpreting different bacteria species in the sample and their interactions as a whole. We show that, compared to baseline representations, our sample representation consistently leads to improved performance for multiple prediction tasks including predicting Irritable Bowel Disease (IBD) and diet patterns. Coupled with a simple ensemble strategy, it produces a highly robust IBD prediction model that generalizes well to microbiome data independently collected from different populations with substantial distribution shift.</p> <p>We visualize the contextualized taxa representations and find that they exhibit meaningful phylum-level structure, despite never exposing the model to such a signal. Finally, we apply an interpretation method to highlight bacterial species that are particularly influential in driving our model's predictions for IBD.</p>

opencc-zeroJun 2024View details →
zenodo40/100

SLDeep: Statement-Level Software Defect Prediction Using Deep-Learning Models on Static Code Features

<p>Software defect prediction (SDP) seeks to estimate fault-prone areas of the code to focus testing activities on more suspicious portions. Consequently, high-quality software is released with less time and effort. The current SDP techniques however work at coarse-grained units, such as a module or a class, putting some burden on the developers to locate the fault. To address this issue, we propose Statement-Level software defect prediction using Deep-learning model (SLDeep). To reify our proposal, we defined a suite of 32 statement-level metrics, such as the number of binary and unary operators used in a statement. Then, we applied as learning model, long short-term memory (LSTM). The significance of SLDeep for intelligent and expert systems is that it demonstrates a novel use of deep-learning models to the solution of a practical problem faced by software developers. We conducted experiments using more than 100,000 C/C++ programs within the Code4Bench. The programs total 2,356,458 lines of code with 292,064 faulty lines. The benchmark comprises diverse set of programs and versions, written by thousands of developers. Therefore, it tends to give a model that can be used for cross-project SDP. In the experiments, our trained model could successfully classify the unseen data with average performance measures 0.945, 0.971, and 0.976 in terms of recall, precision, and accuracy, respectively. These experimental results suggest that SLDeep is effective for statement-level SDP. The impact of this work is twofold. Working at statement-level further alleviates developer&rsquo;s burden in pinpointing the fault locations. Second, cross-project feature of SLDeep helps defect prediction research become more industrially-viable</p> <p>for more information visit&nbsp;<a href="https://github.com/sldeep/SLDeep">https://github.com/sldeep/SLDeep</a></p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Exploring deep learning models for 4D-STEM-DPC data processing

<p>This repository contains scanning transmission electron microscopy data and processing files used in the journal publication&nbsp;<strong>"Exploring deep learning models for 4D-STEM-DPC data processing"</strong>. DOI: <a href="https://doi.org/10.1016/j.ultramic.2024.114058">10.1016/j.ultramic.2024.114058</a></p> <p><strong>Prerequisites</strong></p> <p>The scripts presented below require certain open-source Python packages to run. Library versions used to run the scripts are:</p> <ul> <li>hyperspy 1.7.1</li> <li>pyxem 0.14.2</li> <li>fpd 0.2.5</li> <li>pytorch 1.12.1 (cudatoolkit 11.6.0)</li> <li>jupyterlab 4.0.7</li> </ul> <p><strong>Data files</strong></p> <p>Three zipped folders are included. Two of them contain the training- and inference data for the neural networks, aptly named&nbsp;<em>training_data.zip</em> and&nbsp;<em>inference_data.zip</em>. PyTorch state dictionaries for trained models are included in the&nbsp;<em>models.zip</em> folder.</p> <p><strong>Processing scripts</strong></p> <p>All scripts are included in an IPython notebook format (.ipynb extension). The notebooks&nbsp;<em>Segmentation.ipynb</em> and&nbsp;<em>Regression.ipynb</em> contain the code for training and inference of the segmentation and regression models, respectively. The&nbsp;<em>Training_data_creation.ipynb<strong>&nbsp;</strong></em>notebook contains the code to preprocess the training data for both neural network models. The <em>Standard_algorithms.ipynb</em> notebook has the code for doing center of mass and edge filtering/disc detection algorithms for STEM-DPC processing.</p>

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record