Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
251
datasets available to search
ShareScore release 0.7.1
Dataset results
251 results for “deep learning models”
"An efficient ptychography reconstruction strategy through fine-tuning of large pre-trained deep learning model" train and test data
<ul><li>Model for the article "An efficient ptychography reconstruction strategy through fine-tuning of large pre-trained deep learning model".</li><li>The .pth file is the pre-trained PtyNet-S model and the fine-tuned PtyNet-B model.</li><li>Please contact panxy@ihep.ac.cn if you have any questions.</li></ul>
Dataset associated with "Emulating subglacial hydrology in ice sheet models with deep learning methods" by Verjans and Robel.
<p>See Readme file for descriptions.</p>
Coupling deep learning and physically-based hydrological models for monthly streamflow predictions
<p>Revision in journal Water Resources Research, Manuscript number: <strong><span>2023WR035618R</span></strong></p> <p><strong>Abstract:</strong><strong> </strong>This study proposes a new hybrid model for monthly streamflow predictions by coupling a physically-based distributed hydrological model with a deep learning (DL) model. Specifically, a simplified hydrological model is first developed by optimally selecting grid cells from a distributed hydrological model according to their soil moisture characteristics. <span>It</span> is then driven by bias corrected general circulation model (GCM) <span>prediction</span>s to generate soil moistures for the forecasting months. Finally, model-simulated soil moisture along with other predictors from multiple sources are used as inputs of the DL model to predict future <span>monthly </span>streamflows. The proposed hybrid model, using the simplified Variable Infiltration Capacity (VIC) as the hydrological model and the combination of Convolutional Neural Network and Gated Recurrent Unit (CNN-GRU) as the DL model, is applied to predict 1-, 3-, and 6-month ahead <span>reservoir </span>inflows <span>for the Danjiangkou Reservoir in China. </span>The results show that the hybrid model consistently performs better than VIC and CNN-GRU models with great improvement in Kling‐Gupta efficiency (KGE) values for lead times up to 6 months. <span>Additional tests indicate that hybrid</span> model<span>s based on CNN-GRU </span>outperform <span>those based on</span> <span>LASSO, XGBoost, CNN, and GRU models. Moreover, compared with the distributed hydrological model, the hybrid model</span> greatly reduce<span>s</span> the <span>computation </span>burden of rolling prediction<span>. It also </span>saves decision-makers the time and effort of trying different combinations of predictors<span>, which is indispensable when building DL models. Overall</span>, the new hybrid model <span>demonstrates great potential</span> for monthly streamflow prediction <span>where</span> training data are limited.</p> <p><strong><span>Keywords:</span></strong> <span>monthly streamflow prediction; deep learning; </span><span>physically-based distributed hydrological model; </span><span>VIC model; soil moisture; hybrid model </span></p>
Evaluating the Robustness of Deep-learning Algorithm-selection Models by Evolving Adversarial Instances - Code and Data
<p>This repository contains the code and data for reproducibility of the paper 'Evaluating the Robustness of Deep-learning Algorithm-selection Models by Evolving Adversarial Instances'. </p> <p>The following files are included:</p> <ul> <li>Data.zip : contains the original instances in the datasets;</li> <li>Models.zip : trained Deep Neural Networks models used in the paper;</li> <li>New_instances.zip : generated instances using the approach;</li> <li>Parsed_data.zip : results and statistics of the experiments;</li> <li>script_adversarial_v3.py : Python script used to generate the results</li> </ul>
Enhancing Credit Risk Assessment in Digital Finance through a Hybrid Deep Learning Model Integrated with Blockchain on the Edge of Things F
<p><span>This work proposes a credit risk assessment model using deep learning models such as self-attention generative adversarial networks (SA-GAN) and deep multi-layer perceptron (DMLP). Blockchain is used to improve the security aspects of the model by employing Brakerski-Gentry-Vaikuntanathan (BKV) encryption technique. Further, the proposed system is implemented in Edge-of-things network and communications are enabled via LoRaWAN server.</span></p>
Data and Codes of A Deep Learning-Based Consistency Test for Earth System Models on Heterogeneous Many-Core Systems
<p>These are the supporting information to verify the results in the paper, including input data, model outputs, the postprocessing scripts and the source codes.</p>
Dataset containing images for training and testing deep learning image recognition models
<p>The dataset contains a training and a test subsets of images. Each image belongs to one out of 10 categories of animals</p>
Data Sets for Article: Exploration on Learning Molecular Docking with Deep Learning Models
<p>The MOL2 and CSV file of the clustered compounds from ChemDiv are available in <strong>Additional file 2</strong></p> <p>Docking scores of training set1 and the following traing set2 for each target were saved as csv files and provided in <strong>Additional file 3.</strong></p> <p>The SMILES, MOL2, SDF of DUD-E compounds and PDB of receptors used for validation are provided in <strong>Additional file 4</strong>.</p> <p>The SMILES of 500,000 compounds randomly selected from the ChEMBL database are provided in <strong>Additional file 5</strong></p> <p>The SMILES of compounds with activities from the ChEMBL database for each target are provided in <strong>Additional file 6</strong></p>
Data for paper "Graph Deep Learning Model for Mapping Mineral Prospectivity"
<p>Four prospecting information, namely, NE- and NW- trending faults, Agno Batholithic pluton margins, and porphyry intrusive contacts for mineral prospectivity mapping in Baguio district, Philippines.</p>
Data to publication "The performance of deep generative models for learning joint embeddings of single-cell multi-omics data"
<p>Joint embedding data to publication "The performance of deep generative models for learning joint embeddings of single-cell multi-omics data"</p> <p>Code available at https://github.com/MTreppner/multiomics_dgms</p>
The synthetic and field seismic datasets for "ClinoformNet-1.0: stratigraphic forward modeling and deep learning for seismic clinoform delineation"
<p>This is the synthetic and field siesmic dataset used in manuscript "Three-Dimensional Implicit Structural Modeling Using Convolutional Neural Network". The dimensions of the large-scale and small-scale synthetic seismic datasets are 1600×256 pixels and 900×256 pixels. The field seismic data contains the subsets of the F3 Block, Australia Poseidon, Alaska North Slope seismic data.</p>
Blinded Predictions and Post-hoc Analysis of the Second Solubility Challenge Data: Exploring Training Data and Feature Set Selection for Machine and Deep Learning Models
<p>Training and test datasets and scripts for training models.</p>
Dataset and results for "Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation"
<p>Dataset and results for "Comparing machine learning and deep learning models for probabilistic post-processing of satellite precipitation-driven streamflow simulation"</p> <p>Yuhang Zhang1, Aizhong Ye1*, Phu Nguyen2, Bita Analui2, Soroosh Sorooshian2, Kuolin Hsu2</p> <p>1 State Key Laboratory of Earth Surface Processes and Resource Ecology, Faculty of Geographical Science, Beijing Normal University, Beijing 100875, China.</p> <p>2 Center for Hydrometeorology and Remote Sensing, Department of Civil and Environmental Engineering, University of California, Irvine, Irvine, California, CA 92697, USA.</p> <p>## Dataset </p> <p>Streamflow simulations from one observed precipitation (CMA) and three satellite precipitation products (PDIR, IMERG-F, and GSMaP) for 522 sub-basins.</p> <p>- Q-CMA (streamflow reference)<br> - Q-PDIR (uncorrected)<br> - Q-IMERGF (uncorrected)<br> - Q-GSMAP (uncorrected)</p> <p>### Data structure</p> <p>- Head section (row1-row5)<br> - SubNO: 522 <br> - BeginT: 2003-01-01 00:00 <br> - EndT: 2019-12-31 00:00 <br> - Interval: 1440s (daily)<br> - Revise: 10 (scaling factor to keep int datatype)<br> - Point1 Point2 ... (Subbasin No.)<br> - Data section<br> - 6209 rows, 522 cols</p> <p>## Results</p> <p>Two post-processing model results for test period (2015-1-1 to 2018-12-31).</p> <p>### Data structure</p> <p>- 1462 rows, every row denotes each day from 2015-1-1 to 2018-12-31</p> <p>- 100 columns, every column denotes each quantile from 0.005 to 0.995, total 100 quantiles.</p> <p>### qrf-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p>### lstm-output</p> <p>- pdir (single input)<br> - imergf (single input)<br> - gsmap (single input)<br> - all (multiple inputs)</p> <p> </p>
EMRNA: Accurate RNA structure determination from cryo-EM maps by deep learning and integrated modeling
<p>EMRNA: Accurate RNA structure determination from cryo-EM maps by deep learning and integrated modeling.</p><p>Here stores the input files and output structures of EMRNA and the reproduction result of auto-DRRAFTER.</p>
Whole-body X-ray images of laying hens with keel bone annotations and masks for training deep learning models
<p>This dataset contains whole-body x-ray images of laying hens (n=1051), with the corresponding keel bone annotations and masks. This dataset was basically used to train deep learning models to segment the keel bone from the whole-body x-ray images. But can be used by others for further developments of similar models. This dataset was generated during the research project funded by Svenska Forskningsrådet Formas (2019-02116). The project aimed to develop a digital tool to assess bones of laying hens using x-ray imaging. All images are in JPEG format. All images are named with informative codes indicating bird ID, date and time of x-raying. For instance, in this "001_20230419_0957_PM.Dicom.jpg" image name, "001" stands for the bird ID, "20230419" for the x-raying date, "0957" for the x-raying time, and "PM.Dicom.jpg" indicates that the bird was x-rayed postmortem "PM" with "Dicom" file output and converted to the ".jpg" format. </p> <p>Update on 28/08/2024: This dataset is linked to publication</p> <p>Sallam et al. (2024) Research Note: A deep learning method segments chicken keel bones from whole-body X-ray images,<br>Poultry Science, Volume 103, Issue 11, 2024, 104214, ISSN 0032-5791,<br>https://doi.org/10.1016/j.psj.2024.104214</p>
Integrated Image-based Deep Learning and Language Models for Primary Diabetes Care
<h2><strong>Example data for the paper "Integrated Image-based Deep Learning and Language Models for Primary Diabetes Care"</strong></h2> <p><strong>Example data for the paper "Integrated Image-based Deep Learning and Language Models for Primary Diabetes Care"</strong></p>
Processed data for the manuscript: Promoting Multi-Task Learning as a General Approach for Deep-Learning-based Hydrological Models
<div> <div>Below is a brief overview of the processed data in this repository:</div> <br> <div>- camels_streamflow: This directory contains streamflow data for CAMELS basins covering the period from January 1, 2015, to December 31, 2021. We have not included the original CAMELS dataset, which contains attributes, meteorological forcing, and streamflow data from January 1, 1980, to December 31, 2014, as it can be easily downloaded from the CAMELS website (https://gdex.ucar.edu/dataset/camels.html) and is too large for us to upload to Zenodo.</div> <div>- modiset4camels: This directory includes multiple versions of basin-mean Evapotranspiration (ET) data retrieved from the MOD16A2 data product. The dataset spans from January 1, 2001, to December 31, 2021, with an 8-day temporal resolution.</div> <div>- nldas4camels: This directory contains basin-mean daily meteorological forcing data from the NLDAS-2 dataset, obtained via Google Earth Engine (GEE). The dataset covers the period from January 1, 2001, to December 31, 2021.</div> <div>- smap4camels: This directory features basin-mean Soil Moisture (SSM) data from the NASA-USDA Enhanced SMAP Global Soil Moisture dataset, covering the period from April 2, 2015, to October 3, 2021. The dataset provides SSM measurements at a 5 cm depth. Additionally, we provide basin-mean daily SMAP L4 data spanning from April 1, 2015, to December 31, 2023.</div> </div>
DTU 10MW reference turbine HAWC2 simulations for Model-free estimation of available power with deep learning training
<p>The time series of DTU 10MW HAWC2 model simulations of two channels: hub-height wind speed and produced power. They are generated to train model-free estimation of available power approach, using wind speed and its moving standard deviation as inputs. They include 3-hour length 100Hz simulations of 3 mean wind speeds (7 m/s, 9m/s and 11m/s) as well as 3 levels of turbulence intensity (TI = 7%, 10% and 20%). </p> <p>The dataset and the training algorithm can also be found here: <a href="https://gitlab.windenergy.dtu.dk/tuhf/deep-learning-for-available-power-estimation/tree/master">https://gitlab.windenergy.dtu.dk/tuhf/deep-learning-for-available-power-estimation/tree/master</a></p>
Deep Learning for Black-Box Modeling of Audio Effects
<p>Accompanying audio samples for the paper:</p> <p>Martínez Ramírez M. A., Benetos, E. and Reiss J. D., “Deep Learning for Black-Box Modeling of Audio Effects” submitted to the Applied Sciences: Acoustics and Vibration - Digital Audio Effects special issue, December, 2019.</p> <p>Dry and wet bass and guitar recordings.</p> <p>Dry notes are taken from the IDMT-SMT-Audio-Effects dataset. Author: Michael Stein (Fraunhofer IDMT) https://www.idmt.fraunhofer.de/en/business_units/m2d/smt/audio_effects.html</p> <p>Wet notes are recorded from the limiter and preamplifier of a Universal Audio 6176 Vintage Channel Strip and from the horn and woofer of a 145 Leslie speaker cabinet.</p>
Hcropland30: A hybrid 30-m global cropland map by leveraging global land cover products and Landsat data based on a deep learning model
<p><strong>Hcropland30</strong><strong>:</strong><strong>A 30-m global cropland map by leveraging global land cover products and Landsat data based on a deep learning model</strong></p> <p><strong>***Please note this dataset is undergoing peer review***</strong></p> <p><strong>Version</strong>: <strong>1.0</strong></p> <p><strong>Authors</strong>: Qiong Hu <sup>a, 1</sup>, Zhiwen Cai<sup> b, 1</sup>, Liangzhi You<sup> c, d</sup>, Steffen Fritz<sup> e</sup>, Xinyu Zhang<sup> c</sup>, He Yin<sup> f</sup>, Haodong Wei<sup>c</sup>, Jingya Yang<sup> g</sup>, Zexuan Li<sup> a</sup>, Qiangyi Yu<sup> g</sup>, Hao Wu<sup> a</sup>, Baodong Xu<sup> b *</sup>, Wenbin Wu<sup> g, *</sup></p> <p><em><sup>a</sup></em><em> Key Laboratory for Geographical Process Analysis & Simulation of Hubei Province/College of Urban and Environmental Sciences, Central China Normal University, Wuhan 430079, China</em></p> <p><em><sup>b</sup></em><em> College of Resources and Environment, Huazhong Agricultural University, Wuhan 430070, China</em></p> <p><em><sup>c</sup></em><em> Macro Agriculture Research Institute, College of Plant Science and Technology, Huazhong Agricultural University, Wuhan 430070, China</em></p> <p><em><sup>d </sup></em><em>International Food Policy Research Institute, 1201 I Street, NW, Washington, DC 20005, USA</em></p> <p><em><sup>e </sup></em><em>Novel Data Ecosystems for sustainability Research Group, International Institute for Applied Systems Analysis (IIASA), Schlossplatz 1, Laxenburg A-2361, Austria</em></p> <p><em><sup>f </sup></em><em>Department of Geography, Kent State University, 325 S. Lincoln Street, Kent, OH 44242, USA</em></p> <p><a name="_Hlk166672711"></a><em><sup>g </sup></em><em>State Key Laboratory of Efficient Utilization of Arid and Semi-arid Arable Land in Northern China, the Institute of Agricultural Resources and Regional Planning, Chinese Academy of Agricultural Sciences, Beijing 100081, China</em></p> <p><strong> </strong></p> <p><strong>Introduction</strong></p> <p>We are pleased to introduce a comprehensive global cropland mapping dataset (named Hcropland30) in 2020, meticulously curated to support a wide range of research and analysis applications related to agricultural land and environmental assessment. This dataset encompasses the entire globe, divided into 16,284 grids, each measuring an area of 1°×1°. Hcropland30 was produced by leveraging global land cover products and Landsat data based on a deep learning model. Initially, we established a hierarchal sampling strategy that used the simulated annealing method to identify the representative 1°×1° grids globally and the sparse point-level samples within these selected 1°×1°grids. Subsequently, we employed an ensemble learning technique to expand these sparse point-level samples into the densely pixel-wise labels, creating the area-level 1°×1° cropland labels. These area-level labels were then used to train a U-Net model for predicting global cropland distribution, followed by a comprehensive evaluation of the mapping accuracy.</p> <p> </p> <p><strong>Dataset</strong></p> <p><strong><em><u>1. Hcropland30</u></em></strong><strong>:</strong> A hybrid 30-m global cropland map in 2020</p> <p>****<strong>Data format</strong>: GeoTiff</p> <p>****<strong>Spatial resolution</strong>: 30 m</p> <p>****<strong>Projection</strong>: EPSG: 4326 (WGS84)</p> <p>****<strong>Values</strong>: 1 denotes cropland and 0 denotes non-cropland</p> <p>The dataset has been uploaded in 16,284 tiles. The extent of each tile can be found in the file of “Grids.shp”. Each file is named according to the grid’s Id number. For example, “000015.tif” corresponds to the cropland mapping result for the 15-th 1°×1° grid. This systematic naming convention ensures easy identification and retrieval of the specific grid data.</p> <p><strong><em><u>2. </u></em></strong><strong><em><u>1°×1° </u></em></strong><strong><em><u>Grids</u></em></strong><strong>:</strong> This file contains all 16,284 1°×1° grids used in the dataset. The vector file includes 18 attribute fields, providing comprehensive metadata for each grid. These attributes are essential for users who need detailed information about each grid’s characteristics.</p> <p>****<strong>Data format</strong>: ESRI shapefile</p> <p>****<strong>Projection</strong>: EPSG: 4326 (WGS84)</p> <p>****<strong>Attribute Fields</strong>:</p> <p><strong>Id:</strong> The grid’s ID number.</p> <p><strong>area:</strong> The area of the grid.</p> <p><strong>mode:</strong> Indicates the representative sample grid.</p> <p><strong>climate:</strong> The climate type the grid belongs to.</p> <p><strong>dem: </strong>Average DEM value of the grid.</p> <p><strong>ndvi_s1 to ndvi_s4:</strong> Average NDVI values for four seasons within the grid.</p> <p><strong>esa, esri, fcs30, fromglc, glad, globeland30:</strong> Proportion of cropland pixels of different publicly available cropland products.</p> <p><strong>inconsistent:</strong> Proportion of inconsistent pixels within the grid according to different public cropland products.</p> <p><strong>hcropland30:</strong> Proportion of cropland pixels of our Hcropland30 dataset.</p> <p><strong><em><u>3. Samples</u></em></strong>: The selected representative pixel-level samples, including 32,343 cropland and 67657 non-cropland samples. The category information of each sample was determined based on visual interpretation on Google Earth image and three-year NDVI time series curves from 2019-2021.</p> <p>****<strong>Data format</strong>: ESRI shapefile</p> <p>****<strong>Projection</strong>: EPSG: 4326 (WGS84)</p> <p>****<strong>Attribute Fields</strong>:</p> <p><strong>type:</strong> 1 denotes cropland sample and 0 denotes non-cropland sample.</p> <p><strong>Citation</strong></p> <p>If you use this dataset, please cite the following paper:</p> <p>Hu, Q., Cai, Z., You, L., Fritz, S., Zhang, X., Yin, H., Wei, H., Yang, J., Li, Z., Yu, Q., Wu, H., Xu, B., Wu, W. (2024). Hcropland30: A 30-m global cropland map by leveraging global land cover products and Landsat data based on a deep learning model, Remote Sensing of Environment, submitted.</p> <p><strong>License</strong></p> <p>The data is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0).</p> <p><strong>Disclaimer</strong></p> <p>This dataset is provided as-is, without any warranty, express or implied. The dataset author is not</p> <p>responsible for any errors or omissions in the data, or for any consequences arising from the use</p> <p>of the data.</p> <p><strong>Contact</strong></p> <p>If you have any questions or feedback regarding the dataset, please contact the dataset author</p> <p>Qiong Hu (huqiong@ccnu.edu.cn)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.