Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

59

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

59 results for “Deep-learning”

Learn how ShareScore rates datasets ↗
zenodo48/100

Predictability Limit of the 2021 Pacific Northwest Heatwave from Deep-Learning Sensitivity Analysis

<p>The attached two datasets are the optimized inputs used to analyze predictability limits in the paper Predictability Limit of the 2021 Pacific Northwest Heatwave from Deep-Learning Sensitivity Analysis. Specifically, the datasets correspond to the inputs used to produce the blue (global) and green (regional) loss curves in Figure S2. They are NetCDF files of dimensions batch (1), time (2), latitude (181), longitude (360), pressure levels (13), and may be run as Graphcast model inputs to initiate a forecast at 00 UTC 20 June 2021. Both datasets have been systematically perturbed to reduce the Graphcast model's loss function, which minimizes forecast eror as described in the manuscript. The global input seeks to reduce the loss over the entire globe, while the regional input seeks only to minimize error within the Pacific Northwest (42N to 60N and 130W to 110W). The optimized inputs result in a reduction of the loss by approximately 85% (global) and 93% (regional) when compared to a control Graphcast forecast without perturbations.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Dataset for "Solving deep-learning density-functional theory via variational autoencoder"

<p>The dataset contains the ground state energies, the ground state density profiles, and the external potentials of a 3D single particle system with a Gaussian-like external potential.<br>The number of grid points for each dimension is \(N_g=18\), the linear length of the box is \(L=a_0\) with \(a_0\) the unit of length. The unit of energy is \(E_0=\frac{ \hbar^2}{(m a_0^2)}\).</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>The dataset is zip file of a Python npz file with the following keys:</p> <p>- "density" that corresponds to the ground state density profile.<br>- "potential" is the external potential.<br>-"energy" is the ground state energy.</p> <p><br>The number of instances is 36000.&nbsp;</p> <p>-3D_gaussian.zip -&gt; 3D_gaussian.npz</p> <p>&nbsp; &nbsp; a dictionary with three keys -density, potential, energy-.<br>&nbsp; &nbsp; The dimension of both potential and density is \([N_d,N_g,N_g,N_g]\).<br>&nbsp; &nbsp; The shape of energy is \([N_d]\).<br>&nbsp; &nbsp; \(N_d=36000\)</p> <p>-3D_gaussian_transfer_test_1.npz</p> <p>&nbsp; &nbsp; a dictionary with three keys -density, potential, energy-.<br>&nbsp; &nbsp; The dimension of both potential and density is \([N_d,N_g,N_g,N_g]\).<br>&nbsp; &nbsp; The shape of energy is \([N_d]\).<br>&nbsp; &nbsp; \(N_d=500\)</p> <p>-3D_gaussian_transfer_test_2.npz</p> <p>&nbsp; &nbsp; a dictionary with three keys -density, potential, energy-.<br>&nbsp; &nbsp; The dimension of both potential and density is \([N_d,N_g,N_g,N_g]\).<br>&nbsp; &nbsp; The shape of energy is \([N_d]\).<br>&nbsp; &nbsp; \(N_d=500\)</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Figure 2 in Using deep-learning for automatic identification of images of marine benthic macro-invertebrate bycatch: a proof of concept

Figure 2. – Three images of organisms obtained by cropping images of lots; from left to right: Chalinidae (Porifera), Polyclinidae (Chordata), Hormatidae (Cnidaria).

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 3 in Using deep-learning for automatic identification of images of marine benthic macro-invertebrate bycatch: a proof of concept

Figure 3. – Image of a batch of macro-invertebrate bycatch organisms from Kerguelen Exclusive Economical Zone (Poker 4 survey, 2017), including corals, a crinoïd, an ophiurid, a sea urchin and a brachiopoda; organisms are incomplete and have been quickly spread out over a small plate to take the picture.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 6 in Using deep-learning for automatic identification of images of marine benthic macro-invertebrate bycatch: a proof of concept

Figure 6. – Example of detection and classification obtained with an image including an Ophiuroid, a piece of coral and a sea star with network 2; red squares and annotations have been provided by the computer with no human action.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 5 in Using deep-learning for automatic identification of images of marine benthic macro-invertebrate bycatch: a proof of concept

Figure 5. – Example of detection and classification obtained with an image including Ascidians and a sea star with network 2; red squares and annotations have been provided by the computer with no human action.

opencc-by-4.0Dec 2023View details →
zenodo40/100

SLDeep: Statement-Level Software Defect Prediction Using Deep-Learning Models on Static Code Features

<p>Software defect prediction (SDP) seeks to estimate fault-prone areas of the code to focus testing activities on more suspicious portions. Consequently, high-quality software is released with less time and effort. The current SDP techniques however work at coarse-grained units, such as a module or a class, putting some burden on the developers to locate the fault. To address this issue, we propose Statement-Level software defect prediction using Deep-learning model (SLDeep). To reify our proposal, we defined a suite of 32 statement-level metrics, such as the number of binary and unary operators used in a statement. Then, we applied as learning model, long short-term memory (LSTM). The significance of SLDeep for intelligent and expert systems is that it demonstrates a novel use of deep-learning models to the solution of a practical problem faced by software developers. We conducted experiments using more than 100,000 C/C++ programs within the Code4Bench. The programs total 2,356,458 lines of code with 292,064 faulty lines. The benchmark comprises diverse set of programs and versions, written by thousands of developers. Therefore, it tends to give a model that can be used for cross-project SDP. In the experiments, our trained model could successfully classify the unseen data with average performance measures 0.945, 0.971, and 0.976 in terms of recall, precision, and accuracy, respectively. These experimental results suggest that SLDeep is effective for statement-level SDP. The impact of this work is twofold. Working at statement-level further alleviates developer&rsquo;s burden in pinpointing the fault locations. Second, cross-project feature of SLDeep helps defect prediction research become more industrially-viable</p> <p>for more information visit&nbsp;<a href="https://github.com/sldeep/SLDeep">https://github.com/sldeep/SLDeep</a></p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Dataset for deep-learning in-situ classification of HIV-1 virion morphology

<p>This dataset contains TEM micrographs for HIV-1 virion samples intended for classification and detection as follows:</p> <ol> <li> <p>HIV-1_virion_classification_backbone_dataset.zip : Contains 1806 .tif images of isolated HIV-1 virions extracted and augmented from TEM micrographs. The images are divided into training (1443 images) and validation (363 images) sets and each of these is divided into eccentric, mature, immature labeled folders:</p> <ul> <li> <p>HIV-1_virion_classification_backbone_dataset/</p> <ul> <li> <p>train/</p> <ul> <li> <p>eccentric/</p> </li> <li> <p>immature/</p> </li> <li> <p>mature/</p> </li> </ul> </li> <li> <p>val/</p> <ul> <li> <p>eccentric/</p> </li> <li> <p>immature/</p> </li> <li> <p>mature/</p> </li> </ul> </li> </ul> </li> </ul> </li> <li> <p>HIV-1_rcnn_dataset_full.zip : Contains 59 .tif TEM micrographs of HIV-1 samples as well as a matching .csv file recording the attributes of each viral instance and coordinates of the rectangular region that contains it:</p> </li> </ol> <ul> <li> <p>region_data_&lt;image_id&gt;.csv:</p> <ul> <li> <p>filename: Name of the image this csv refers to. (Ex: 0131001.png)</p> </li> <li> <p>file_size: Size (bytes) of the image this csv refers to. (Ex: 11755590)</p> </li> <li> <p>file_attributes: Specific attributes of the micrograph. (Ex: None)</p> </li> <li> <p>region_count: Number of viral instances detected in the micrograph. (Ex: 39)</p> </li> <li> <p>region_id: ID of a specific viral region. (Ex: 1)</p> </li> <li> <p>region_shape_attributes: Coordinates of the bounding box of &lt;region_id&gt; that contains a virion. (Ex: {&quot;name&quot;:&quot;rect&quot;,&quot;x&quot;:1022,&quot;y&quot;:357,&quot;width&quot;:225,&quot;height&quot;:228})</p> </li> <li> <p>region_attributes: Classification of the virion enclosed in this region (eccentric/mature/immature). (Ex: {&quot;particle_class&quot;:&quot;mature&quot;})</p> </li> </ul> </li> </ul> <p>The images are divided into training (46 images) and validation (13 images) sets and each of these contains folder for each micrograph:</p> <ul> <li> <p>HIV-1_rcnn_dataset_full/</p> <ul> <li> <p>train/</p> <ul> <li> <p>0131001/</p> <ul> <li> <p>0131001.png</p> </li> <li> <p>region_data_0131001.csv</p> </li> </ul> </li> <li> <p>0131004/</p> <ul> <li> <p>0131004.png</p> </li> <li> <p>region_data_0131004.csv</p> </li> </ul> </li> <li> <p>&hellip;</p> </li> </ul> </li> <li> <p>val/</p> <ul> <li> <p>0131002/</p> <ul> <li> <p>0131002.png</p> </li> <li> <p>region_data_0131002.csv</p> </li> </ul> </li> <li> <p>0131003/</p> <ul> <li> <p>0131003.png</p> </li> <li> <p>region_data_0131003.csv</p> </li> </ul> </li> <li> <p>&hellip;</p> </li> </ul> </li> </ul> </li> </ul> <p>The first dataset (HIV-1_virion_classification_backbone_dataset.zip) is intended for viral classification algorithms while the second dataset (HIV-1_rcnn_dataset_full) is intended for detection and segmentation algorithms (for example RCNN).</p> <p>Applications of this dataset as well as source code can be found at <a href="https://github.com/Perilla-lab/TEMNet">https://github.com/Perilla-lab/TEMNet</a> .</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

Dataset for: A Deep-Learning Technique to Locate Cryptographic Operations in Side-Channel Traces

<p>This dataset is part of "A Deep-Learning Technique to Locate Cryptographic Operations in Side-Channel Traces" available <a href="https://www.arxiv.org/abs/2402.19037" target="_blank" rel="noopener">online</a>.</p> <p>The source code for testing the dataset is available on <a href="https://github.com/hardware-fab/DL-to-locate-COs-for-SCA">GitHub</a>.</p> <p>The dataset is organized as follows:</p> <ul> <li><strong>\training</strong>: contains three subsets, i.e., train, valid, and test. &nbsp;<br>&nbsp;&nbsp;Each subset consists of two .npy files: <ul> <li><em>_set</em>: it contains the side-channel traces that are preprocessed accordingly.</li> <li>&nbsp;<em>_labels</em>: itcontains the target labels for training the CNN, labeling each data as <em>cipher start</em>, <em>cipher rest</em>, or <em>noise</em>.</li> </ul> </li> <li><strong>\inference</strong>: contains two files as a demo of the inference pipeline.<br>&nbsp; &nbsp; One file is the is the side-channel trace containing an undefined number of AES encryptions. The other file is a list of plaintexts matching the AES encryptions to test a CPA attack.</li> </ul> <p><strong>Cite:</strong></p> <blockquote> <pre><code>@INPROCEEDINGS{10546758, author={Chiari, Giuseppe and Galli, Davide and Lattari, Francesco and Matteucci, Matteo and Zoni, Davide}, booktitle={2024 Design, Automation &amp; Test in Europe Conference &amp; Exhibition (DATE)}, title={A Deep- Learning Technique to Locate Cryptographic Operations in Side-Channel Traces}, year={2024}, pages={1-6}, doi={10.23919/DATE58400.2024.10546758}}</code></pre> </blockquote> <p>This repository is protected by copyright and licensed under the <a href="https://creativecommons.org/licenses/by/4.0/legalcode">Creative Commons Attribution 4.0 International</a> license.</p> <p>&copy; 2024 hardware-fab</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Dataset for: Hound: Locating Cryptographic Primitives in Desynchronized Side-Channel Traces Using Deep-Learning

<p>This dataset is part of "Hound: Locating Cryptographic Primitives in Desynchronized Side-Channel Traces Using Deep-Learning" [1] available&nbsp;<a href="https://arxiv.org/pdf/2408.06296">online</a>.</p> <p>The source code for testing the dataset is available on <a href="https://github.com/hardware-fab/Hound">GitHub</a>.</p> <p>The dataset is organized as follows:</p> <ul> <li><strong>/training</strong>: Contains three subsets: <em>train</em>, <em>valid</em>, and <em>test</em>. Each subset consists of two <em>.npy</em>&nbsp;files: <ul> <li><em><strong>_set</strong></em>: Contains the preprocessed side-channel traces.</li> <li><strong><em>_labels</em></strong>: Contains the target labels for training the CNN, labeling each data as `CP start`, `CP spare`, or `noise`.</li> </ul> </li> <li><strong>/inference</strong>: Contains files for two demos: consecutive AES executions and AES executions interleaved with noisy applications. Each demo consists of two <em>.npy</em>&nbsp;files: <ul> <li><strong>aes_</strong>: Contains the side-channel traces to input into Hound.</li> <li><strong>gt_</strong>: Contains the ground truth for checking the correctness of Hound segmentation.</li> </ul> </li> </ul> <p>This repository is protected by copyright and licensed under the <a href="https://creativecommons.org/licenses/by/4.0/legalcode">Creative Commons Attribution 4.0 International</a> license.</p> <p>&copy; 2024 hardware-fab</p> <blockquote> <p>[1] D. Galli, G. Chiari and D. Zoni, "Hound: Locating Cryptographic Primitives in Desynchronized Side-Channel Traces using Deep-Learning," 2024 IEEE 42nd International Conference on Computer Design (ICCD), Milan, Italy, 2024, pp. 114-121, doi: 10.1109/ICCD63220.2024.00027.</p> </blockquote>

opencc-by-4.0Nov 2024View details →
zenodo36/100

"Toy Data Set" referenced in the article "A deep-learning based analysis framework for ultra-high throughput screening time-series data" (https://doi.org/10.1101/2024.08.22.609110)

<p>This data set, referenced as "toy data set" in the article "A deep-learning based analysis framework for ultra-high throughput screening time-series data" (<a href="Lint-to-article">https://doi.org/10.1101/2024.08.22.609110</a>), mimics a high-throughput screening data set. To demonstrate the application of our analysis framework described in the main article this toy data set was generated. It contains in total 1536000 individual transient signals, splitted in 5 batches of each 200 plates in 1536-well plate format. Five distinct signal classes were used to resemble typical shapes encountered in biological experiments. Fequency of occurrences for each class is reported in the main article.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

A Deep-Learning Approach for Visual Detection of an AUV Docking Station - Dataset

<div>This dataset was used to train the models from the paper "A Deep-Learning Approach for Visual Detection of an AUV Docking Station" published by Ahmad et al. at Oceans 2024 Conference in Halifax.</div> <div>&nbsp;</div> <div># Dataset</div> <div>&nbsp;</div> <div>The dataset contains:</div> <div>&nbsp;</div> <div>1. images from Abisko Lake in Sweden[1]</div> <div>2. images recorded in the Maritime Hall basin at DFKI</div> <div>&nbsp;</div> <div># File contents</div> <div>Each of these datasets are put into seperate directories. The images were annotated using CVAT[2].</div> <div>&nbsp;</div> <div>The dataset has been exported into the following formats:</div> <div>&nbsp;</div> <div>1. YOLO</div> <div>2. PascalVOC</div> <div>3. COCO</div> <div>&nbsp;</div> <div>The exported datasets does not contain raw images, rather they are places into a seperate zip folder.</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;</div> <div># References</div> <div>[1]: <a href="https://zenodo.org/records/7035132">https://zenodo.org/record/7035132#.ZDfKE5FBzJU</a></div> <div>[2]: https://www.cvat.ai/</div>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Glacier catalogue for IGM physics-informed deep-learning emulator pretraining

<p>This dataset was created with the iceflow glacier model CfsFlow to generate glacier extent and retreat in the Alps and New zealand with the goal to generate realistic and diverse glacier states for pretraining the physics-informed deep-learning emulator of IGM (https://github.com/jouvetg/igm).</p> <p>The data consists of distributed surface topography (usurf) and ice thickness (thk) of 8 snapshots of 37 glaciers in different stages (advance and retreat). The data is organized glacier-wise: each folder corresponds to one glacier, which contains a unique NetCDF file with 2D distributed raster data of surface elevation and ice thickness.</p>

opencc-by-4.0Sep 2023View details →
ClinicalTrials.gov36/100

Deep-Learning Image Reconstruction in CCTA

ClinicalTrials.gov study NCT03980470. IPD Sharing: NO. Countries: 1. Publications: 5.

closedIPD-NOFeb 2026View details →
dryad36/100

Data from: Aerobatic maneuvers in insect-scale flapping-wing aerial robots via deep-learned robust tube model predictive control

Open the record for dataset details and reuse information.

publicNov 2025View details →
dryad36/100

3D photogrammetry and deep-learning deliver accurate estimates of epibenthic biomass

Open the record for dataset details and reuse information.

publicMar 2024View details →
zenodo32/100

Evaluating the Robustness of Deep-learning Algorithm-selection Models by Evolving Adversarial Instances - Code and Data

<p>This repository contains the code and data for reproducibility of the paper 'Evaluating the Robustness of Deep-learning Algorithm-selection Models by Evolving Adversarial Instances'.&nbsp;</p> <p>The following files are included:</p> <ul> <li>Data.zip : contains the original instances in the datasets;</li> <li>Models.zip : trained Deep Neural Networks models used in the paper;</li> <li>New_instances.zip : generated instances using the approach;</li> <li>Parsed_data.zip : results and statistics of the experiments;</li> <li>script_adversarial_v3.py : Python script used to generate the results</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo32/100

Dataset for deep-learning density functional theory Hamiltonian for efficient ab initio electronic-structure calculation

<p>Dataset files of atomic structures and Hamiltonian matrices of graphene, MoS<sub>2</sub>, bilayer graphene&nbsp;and bilayer bismuthene.</p> <p>Please note that the DFT results in this dataset were calculated using OpenMX. This means that if you want to use a DeepH model trained on this dataset to calculate properties, you need to use the&nbsp;<a href="https://github.com/mzjb/overlap-only-OpenMX">overlap calculated using OpenMX</a>. The orbital information required for overlap calculations can be found in the&nbsp;<a href="https://www.nature.com/articles/s43588-022-00265-6">paper</a>.</p>

opencc-by-4.0May 2022View details →
zenodo32/100

Dataset for "Generalizing deep-learning electronic structure calculation to plane-wave basis"

<div> <div>This is the dataset for the paper "Generalizing deep-learning electronic structure calculation to plane-wave basis". For details, please see README.md.</div> </div>

opencc-by-4.0Aug 2024View details →
zenodo32/100

DeepRNA-Reg: A Deep-Learning Based Approach for Comparative Analysis of CLIP Experiments

<p>GEO series accession information for datasets associated with DeepRNA-Reg</p>

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record