Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,273

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,273 results for “Deep learning”

Learn how ShareScore rates datasets ↗
zenodo52/100

Experimental data for "Deep Learning Methods for Colloidal Silver Nanoparticle Concentration and Size Distribution Determination from UV-Vis Extinction Spectra"

<p>Testing data (experimental data) for neural networks published in preprint https://doi.org/10.48550/arXiv.2404.10891</p> <p>The UV-VIS-NIR spectral data was also used in the dissertation of Nadzeya Khinevch, titled "Two-dimensional structures of nanoparticles for elements of surface-enhanced Raman scattering substrates".</p> <p>Emails of the corresponding authors:</p> <p>Tomas Klinavičius tomas.klinavicius@ktu.lt</p> <p>Tomas Tamulevičius tomas.tamulevicius@ktu.lt</p>

opencc-by-4.0Apr 2024View details →
edi52/100

Estimation of Abundance and Distribution of Salt Marsh Plants from Images Using Deep Learning

Recent advances in computer vision and machine learning, most notably deep convolutional neural networks (CNNs), are exploited to identify and localize various plant species in salt marsh images. Three different approaches are explored that provide estimations of abundance and spatial distribution at varying levels of granularity in terms of spatial resolution. In the coarsest-grained approach, CNNs are tasked with identifying which of six plant species are present/absent in large patches within the salt marsh images. CNNs with diverse topological properties and attention mechanisms are shown capable of providing accurate estimations with > 90% precision and recall in the case of the more abundant plant species whereas the performance of the CNNs is observed to decline in the case of less common plant species. Estimation of percent cover of each plant species is performed at a finer spatial resolution, where smaller image patches are extracted and the CNNs tasked with identifying the plant species or substrate at the center of the image patch. In an ecological setting, several image patches (~100) are extracted and classified using this approach to estimate the percent cover of the various plant species in the image. For the percent cover estimation task, the CNNs are observed to exhibit a performance profile similar to that for the presence/absence estimation task, but with an ~ 5–10% reduction in precision and recall. Finally, estimation of the spatial distribution of the various plant species is performed via semantic segmentation of the input images at the finest level of granularity in terms of spatial resolution. The Deeplab-V3 semantic segmentation architecture is observed to provide very accurate estimations for abundant plant species; however, a significant degradation in performance is observed in the case of less abundant plant species and, in extreme cases, rare plant classes are seen to be ignored entirely. Overall, a clear trade-off is observed between

openCC (other)Jan 2020View details →
zenodo48/100

The DNNLikelihood: enhancing likelihood distribution with Deep Learning

<p>Datasets and trained models&nbsp;corresponding to version 2 of <a href="https://arxiv.org/abs/1911.03305">arXiv:1911.03305</a> and complementing&nbsp;the code on&nbsp;<a href="https://github.com/riccardotorre/DNNLikelihood/releases/tag/1911.03305v2">GitHhub</a>.</p> <p>Notice that the code on GitHub includes scripts to automatically download these data.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Bioactivity deep learning for structure-free compound-protein interaction

<p>CPI2M data for "<strong>Bioactivity deep learning for structure-free compound-protein interaction</strong>".</p> <p>CPI2M_main_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_Kd.csv: Bioactivity data with <strong>pKd</strong> activity type. Used for model training and internal validation.</p> <p>CPI2M_main_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_few_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for external validation.</p> <p>CPI2M_few_Kd.csv: Bioactivity data with <strong>pKd </strong>activity type. Used for external validation.</p> <p>CPI2M_few_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for external validation.</p> <p>CPI2M_few_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for external validation.</p> <p>potency.csv: BIoactivity data with <strong>pPotency </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>percentage.csv: BIoactivity data with <strong>Percentage Inhibition </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>Protein_pretrained_feat.zip: pre-calculated protein feature files with UniProt ID naming. <strong>Should be unzipped</strong> before start model training with CPI2M data.</p> <p>&nbsp;</p> <p>For each .csv data, columns include "<strong>smiles</strong>" (ligand SMILES), "<strong>exp_mean</strong>" (nM bioactivity), "<strong>y</strong>" (neg.log nM, final label), "<strong>cliff_mol</strong>" (whether activity cliff or not), "<strong>split</strong>" (splitting label by activity cliff), "<strong>Uniprot_id</strong>" (UniProt ID for protein), "<strong>Sequence</strong>" (wildtype sequence for protein), and "type_id" (bioactivity type token, pKi =0, pKd=1, pEC50=2, pIC50=3).</p> <p>&nbsp;</p> <p>Please find the project code at https://github.com/gu-yaowen/GGAP-CPI</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior

<p>The datasets used in the paper "Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior". A detailed description of these datasets is available at https://github.com/jaydu1/VITAE/tree/master/data.</p>

opencc-by-4.0Dec 2020View details →
zenodo48/100

LigPCDS: Labeled Dataset of X-ray Protein Ligand Images in 3D Point Cloud and Validated Deep Learning Models

<p>The difference electron density from X-ray protein crystallography was used to create the first dataset of labeled ligand images in 3D point clouds, named <strong>LigPCDS</strong>. The dataset contain 244,226 entries of free organic ligands containing 3D representations labeled with two major labeling approaches: SP-based and AtomSymbol-based.</p> <p>&nbsp;</p> <p>The data from free organic molecules (non-covalent ligands) was retrieved from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) in december 2019 with resolutions ranging from 1.5 to 2.2 &Aring;. The ligand images (blobs) were interpolated from their calculated difference electron density map in a 3D grid-like bounding box, around their atomic positions, and stored in point clouds. These ligand grid representations were further processed to retrive the final ligands representation in 3D point clouds using a mask of the shape of the ligand. A grid spacing of 0.5 &Aring; gave the best results. The density value of the grid points was used as feature. The labeling approach used the structure of the ligands to propose vocabularies of chemical classes based on the chemical atoms themselves and their cyclic substructures. These structure annotations were applied pointwise to the ligand 3D representations using an atomic sphere model. Four proposed vocabularies were validated by successfully training good performance deep learning models for the semantic segmentation of a stratified dataset from LigPCDS, using 78902 entries.</p> <p>The four validated deep learning models are: (i) the LigandRegion, composed by generic atoms of any type; (ii) the AtomCycle, composed by generic atoms outside cycles and generic cycles; (iii) the AtomC347CA56, composed by generic atoms outside cycles, not aromatic cycles of size 3 to 7 and aromatic cycles of size 5 and 6; and (iv) the AtomSymbolGroups, composed by the atoms symbols with groupings. The mean accuracy of these models in their cross-validation was between 49.7% <span lang="EN-GB">[-19.4,20.</span><span lang="EN-GB">2]</span> and 77.4% <span lang="EN-GB">[-11.7,12.1]</span> in terms of Intersection over Union (mIoU) metric and between 62.4% <span lang="EN-GB">[-18.8,19.</span><span lang="EN-GB">7]</span> and 87.0% <span lang="EN-GB">[-8.4,8.8]</span> in F1-score (mF1), confidence interval between squared brackets. The models i, ii and iii and the used labeled representations in 3D point cloud are contained in the SP-based record; and model iv and its used labeled representations are contained in the AtomSymbol-based record.</p> <p>The dataset and validated models may be used to tackle problems regarding known and unknown ligand building to drug discovery and fragment screening pipelines.&nbsp;</p> <p>The code used to create and validated the LigPCDS is available at the following repository: https://github.com/danielatrivella/np3_ligand</p> <p>This repository also contains the NP&sup3; Blob Label application for ligand building using the validated deep learning models from LigPCDS.</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study

<p>Efficiency is essential to support responsiveness w.r.t. ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code that supports symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development tends to produce DL code that is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, less error-prone imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. While hybrid approaches aim for the &quot;best of both worlds,&quot; the challenges in applying them in the real world are largely unknown. We conduct a data-driven analysis of challenges&mdash;and resultant bugs&mdash;involved in writing reliable yet performant imperative DL code by studying 250 open-source projects, consisting of 19.7 MLOC, along with 470 and 446 manually examined code patches and bug reports, respectively. The results indicate that hybridization: (i) is prone to API misuse, (ii) can result in performance degradation&mdash;the opposite of its intention, and (iii) has limited application due to execution mode incompatibility. We put forth several recommendations, best practices, and anti-patterns for effectively hybridizing imperative DL code, potentially benefiting DL practitioners, API designers, tool developers, and educators.</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

Sentinel2GlobalLULC: A dataset of Sentinel-2 georeferenced RGB imagery annotated for global land use/land cover mapping with deep learning (License CC BY 4.0)

<p>Sentinel2GlobalLULC is a deep learning-ready dataset of RGB images from the Sentinel-2 satellites designed for global land use and land cover (LULC) mapping. Sentinel2GlobalLULC v2.1&nbsp;contains 194,877 images in GeoTiff and JPEG format corresponding to 29 broad LULC classes. Each image has 224 x 224 pixels at 10 m spatial resolution and was produced by assigning the 25th percentile of all available observations in the Sentinel-2 collection between June 2015 and October 2020 in order to remove atmospheric effects (i.e., clouds, aerosols, shadows, snow, etc.). A spatial purity value was assigned to each image based on the consensus across 15 different global LULC products available in Google Earth Engine (GEE).&nbsp;</p> <p>&nbsp;</p> <p>Our dataset is structured into 3 main zip-compressed folders, an Excel file with a dictionary for class names and descriptive statistics per LULC class, and a python script to convert RGB GeoTiff images into JPEG format. The first folder called &quot;Sentinel2LULC_GeoTiff.zip&quot;&nbsp;contains 29 zip-compressed subfolders where each one corresponds to a specific LULC class with hundreds to thousands of GeoTiff Sentinel-2 RGB images. The second folder called &quot;Sentinel2LULC_JPEG.zip&quot; contains 29 zip-compressed subfolders with a JPEG formatted version of the same images provided in the first main folder. The third folder called &quot;Sentinel2LULC_CSV.zip&quot; includes 29 zip-compressed CSV files with as many rows as provided images and with 12&nbsp;columns containing the following metadata (this same metadata is provided in the image filenames):&nbsp;</p> <ul> <li>Land Cover Class ID: is the identification number of each LULC class</li> <li>Land Cover Class Short Name: is the short name of each LULC class</li> <li>Image ID: is the identification number of each image within its corresponding LULC class&nbsp;</li> <li>Pixel purity Value: is the spatial purity of each pixel for its corresponding LULC class calculated as the spatial consensus across up to 15 land-cover products&nbsp;</li> <li>GHM Value: is the spatial average of the Global Human Modification index (gHM) for each image</li> <li>Latitude: is the latitude of the center point of each image</li> <li>Longitude: is the longitude of the center point of each image</li> <li>Country Code: is the Alpha-2 country code of each image as described in the ISO 3166 international standard. To understand the country codes, we recommend the user to visit the following website where they present the Alpha-2 code for each country as described in the ISO 3166 international standard:https: //www.iban.com/country-codes</li> <li>Administrative Department Level1: is the administrative level 1 name to which each image belongs</li> <li>Administrative Department Level2: is the administrative level 2 name to which each image belongs</li> <li>Locality: is the name of the locality to which each image belongs</li> <li>Number of S2 images : is&nbsp;the number of found instances in the corresponding Sentinel-2 image collection between June 2015 and October 2020, when compositing&nbsp;and exporting&nbsp;its corresponding&nbsp;image tile</li> </ul> <p>For seven LULC classes, we could not export from GEE all images that fulfilled a spatial purity of 100% since there were millions of them. In this case, we exported a stratified random sample of 14,000 images and provided an additional CSV file with the images actually contained in our dataset. That is, for these seven LULC classes, we provide these 2 CSV files:</p> <ul> <li>A CSV file that contains all exported images for this class&nbsp;</li> <li>A CSV file that contains all images available for this class at spatial purity of 100%, both the ones exported and the ones not exported, in case the user wants to export them. These CSV filenames end with &quot;including_non_downloaded_images&quot;.</li> </ul> <p>To clearly state the geographical coverage of images available in this dataset,&nbsp; we&nbsp;included in the version v2.1, &nbsp;a compressed folder called &quot;Geographic_Representativeness.zip&quot;. This zip-compressed folder&nbsp;contains a csv file&nbsp;for each LULC class that provides the complete list of countries represented in that class. Each csv file has two columns, the first one gives the country code and the second one gives the number of images provided in that country for that LULC class. In addition to these 29 csv files, we provided another csv file that maps each ISO Alpha-2 country code to its original full country name.</p> <p>&copy;&nbsp;<a href="https://doi.org/10.5281/zenodo.5055632">Sentinel2GlobalLULC Dataset&nbsp;</a>by&nbsp;&nbsp;Yassir Benhammou, Domingo Alcaraz-Segura, Emilio Guirado, Rohaifa Khaldi, Boujem&acirc;a Achchab, Francisco Herrera &amp; Siham Tabik&nbsp;is marked with Attribution 4.0 International&nbsp;(CC-BY 4.0)</p>

opencc-by-4.0Jul 2022View details →
zenodo48/100

Phase Object Reconstruction for 4D-STEM using Deep Learning, (4D-STEM Example Data)

<p><strong>Overview </strong></p> <p>This repository contains 2 example 4D-STEM datasets format from the paper <a href="https://arxiv.org/abs/2202.12611">&quot;Phase Object Reconstruction for 4D-STEM using Deep Learning&quot;</a>. The data was written to hdf5 for compatibility with the python programming language. When reading from these files consider possibly different storage conventions (Row major vs. column major format). Data may need to be transposed accordingly.</p> <p>&nbsp;</p> <p><strong>Parameters</strong></p> <p>The twisted bilayer graphene dataset is simulated. The smaller file is an experimental SrTiO<sub>3</sub> dataset.</p> <table> <thead> <tr> <th scope="row">&nbsp;</th> <th scope="col">Graphene</th> <th scope="col">STO</th> </tr> </thead> <tbody> <tr> <th scope="row">E0</th> <td>200kV</td> <td>300kV</td> </tr> <tr> <th scope="row">Apeture</th> <td>25 mrad</td> <td>20 mrad</td> </tr> <tr> <th scope="row">Detector Size</th> <td>2.5 &Aring;<sup>-1</sup></td> <td>1.6671 &Aring;<sup>-1</sup></td> </tr> <tr> <th scope="row">Dimensions</th> <td>101x101x128x128</td> <td>60x60x64x64</td> </tr> <tr> <th scope="row">Step Size</th> <td>0.2</td> <td>0.1818</td> </tr> </tbody> </table> <p><br> &nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

Supplemental Figures for "On the comparative utility of entropic learning versus deep learning for long-range ENSO prediction"

<p>Supplemental figures for the paper "On the comparative utility of entropic learning versus deep learning for long-range ENSO prediction".</p>

opencc-by-4.0Jan 2024View details →
zenodo48/100

Deep learning-based earthquake catalog of the 2022 MW 6.9 Chihshang, Taiwan, earthquake sequence

<p>On 18 September 2022, the MW 6.9 Chihshang earthquake struck the southern Longitudinal Valley, Taiwan. We use SeisBlue, a deep-learning platform/package, to extract the two-month earthquake sequence from September to October 2022, including the MW 6.5 Guanshan foreshock, the MW 6.9 mainshock, over 14,000 aftershocks, and 866 focal mechanisms from two sets of broadband networks. For more details, please refer to our research article published at TAO (Sun et al., 2024; https://doi.org/10.1007/s44195-024-00063-9). The refined SeisBlue earthquake, FMS, and 20-year M3+ relocated CWA earthquake catalogs obtained in this study are listed here.</p> <ol> <li>The refined, deep-learning-based earthquake catalog of the 2022 Mw 6.9 Chihshang, Taiwan, earthquake sequence contains 5,151 seismic events with event time, location and error information, local and moment magnitudes, and hypoDD location.&nbsp;</li> <li>The FMS (focal mechanism solution) catalog is obtained by the P-wave polarities of 14 broadband stations and the FPFIT program (Reasenberg &amp; Oppenheimer, 1985). 865 out of 1629 FMSs with at least six readings of P-wave polarity, F-fator <span>&le; </span>0.1 (F <span>&lt; </span>0.5 for a good fit), and errors of strike, dip, and rake are all <span>&lt; </span>20<span>&deg;</span>, respectively, are listed in the attached FMS catalog.</li> <li>The 2001-2020 3D-hypoDD-relocated M3+ CWA earthquake catalog: We applied the HypoDD program (Waldhauser &amp; Ellsworth, 2000) to the CWA (Central Weather Administration (CWA, Taiwan), 2012) catalog with P- and S-wave arrivals and obtained 5862 M3+ events between 2001 and 2020 for eastern Taiwan. The 3D velocity models used for this catalog are the local models from Kuo-Chen et al. (2012).</li> </ol> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo48/100

Predictability Limit of the 2021 Pacific Northwest Heatwave from Deep-Learning Sensitivity Analysis

<p>The attached two datasets are the optimized inputs used to analyze predictability limits in the paper Predictability Limit of the 2021 Pacific Northwest Heatwave from Deep-Learning Sensitivity Analysis. Specifically, the datasets correspond to the inputs used to produce the blue (global) and green (regional) loss curves in Figure S2. They are NetCDF files of dimensions batch (1), time (2), latitude (181), longitude (360), pressure levels (13), and may be run as Graphcast model inputs to initiate a forecast at 00 UTC 20 June 2021. Both datasets have been systematically perturbed to reduce the Graphcast model's loss function, which minimizes forecast eror as described in the manuscript. The global input seeks to reduce the loss over the entire globe, while the regional input seeks only to minimize error within the Pacific Northwest (42N to 60N and 130W to 110W). The optimized inputs result in a reduction of the loss by approximately 85% (global) and 93% (regional) when compared to a control Graphcast forecast without perturbations.</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Navigating deep learning strategies for large-area land cover mapping using very-high-resolution imagery in Senegal: Validation Data

<p><span><span>R</span><span>apid</span><span> advances in deep learning</span><span> for</span> <span>land cover </span><span>classification of </span><span>trees, shrubs and </span><span>very small</span> <span>agricultur</span><span>al</span> <span>fields</span> <span>using</span> <span>very high</span><span>-</span><span>resolution satellite </span><span>data </span><span>(&lt; 2 m</span><span>)</span><span>,</span><span> has tremendous potential</span> <span>for resolving </span><span>current</span><span> challenges </span><span>in </span><span>quantifying</span> <span>land cover </span><span>change </span><span>in</span> <span>sub-</span><span>Saharan</span> <span>African (SSA</span><span>)</span><span>,</span> <span>due to</span> <span>growing </span><span>demand for food resources</span><span>.</span> <span>We</span> <span>conducted experiments </span><span>with</span><span> different training strategies for scaling up </span><span>UNet</span> <span>convolutional neural network </span><span>models for regional land cover mapping with multispectral </span><span>WorldView</span><span> (WV</span><span>)</span><span>-2 and &ndash;3,</span><span> imagery</span><span> in</span><span> three distinct regions of Senegal </span><span>which</span> <span>has</span><span> complex </span><span>seasonal wet/dry conditions and </span><span>cropland-savanna mosaics.&nbsp;</span></span></p> <p>The validation exercise of this research consisted in validating more than 70,000 km<sup>2</sup> across Senegal. The infrastructure was setup in the NASA SMCE system with a total of twelve George Mason University (GMU) students participating as operators. These operators validated more than 59 WV-2 and -3 images, each consisting of 200 stratified points in 5,000 x 5,000-pixel images. This effort resulted in a total of ~35,000 aggregated observations that are available through the eo-validation API for public consumption. Each validation point from this dataset has three individual observations.</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution

<p>Efficiency is essential to support ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code---supporting symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. Though hybrid approaches aim for the "best of both worlds," using them effectively requires subtle considerations. Our key insight is that, while DL programs typically execute sequentially, hybridizing imperative DL code resembles parallelizing sequential code in traditional systems. Inspired by this, we present an automated refactoring approach that assists developers in determining which otherwise eagerly-executed imperative DL functions could be effectively and efficiently executed as graphs. The approach features novel static imperative tensor and side-effect analyses for Python. Due to its inherent dynamism, analyzing Python may be unsound; however, the conservative approach leverages a speculative (keyword-based) analysis for resolving difficult cases that informs developers of any assumptions made. The approach is: (i) implemented as a plug-in to the PyDev Eclipse IDE that integrates the WALA Ariadne analysis framework and (ii) evaluated on nineteen DL projects consisting of 132 KLOC. The results show that 326 of 766 candidate functions (42.56%) were refactorable, and an average relative speedup of 2.16x on performance tests was observed with negligible differences in model accuracy. The results indicate that the approach is useful in optimizing imperative DL code to its full potential.</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Phase unwrapping using Deep Learning in Holographic Tomography - dataset.

<p>This dataset contains two types of data: phase images and trained model files.</p> <ol> <li> <p><strong>Real phase images</strong> - these phase images are contained with the files named with the prefix &quot;real_&quot;. The type of the data files is &quot;<strong>.npz</strong>&quot;, to be loaded with NumPy (<em>np.load()</em>), as a dictionary. The data is stored within the key [&quot;<em>arr_0</em>&quot;]. The images depict cells [1], organoids [2], phantoms [3-4] and regular 3D printed structures with high scattering properties [5]. The images have been augmented in order to expand the volume of the training dataset. All images are of shape (<em>256,256,1</em>). It is a big dataset containing 27,189 images of each type for training the unwrapping model:</p> <ul> <li> <p>unwrapped - continuous phase distribution (<em>float32</em>)</p> </li> <li> <p>wrapped - phase wrapped into mod2<span class="math-tex">\(\pi\)</span> (<em>float32</em>)</p> </li> <li> <p>wrapcount - wrap count phase maps coded in the integer form (0,1,2...) (<em>uint8</em>)</p> </li> </ul> </li> <li> <p><strong>Synthetic phase images</strong> - phase images in these files were generated algorithmically in the MATLAB programming language. The files containing this dataset have a prefix &quot;synthetic_&quot;.&nbsp;The type of the data files is &quot;<strong>.npz</strong>&quot;, to be loaded with NumPy (<em>np.load()</em>), as a dictionary. The data is stored within the key [&quot;<em>arr_0</em>&quot;]. Phase images contained in the synthetic dataset can be split into 3 types by their type: spherical distribution, simulated cells w/ spherical background and simulated cells w/ introduced linear tilt.&nbsp;All images are of shape (<em>256,256,1</em>). This dataset contains 10,000 images of each type for training the unwrapping and denoising models:</p> <ul> <li> <p>unwrapped - continuous phase distribution (<em>float32</em>)</p> </li> <li> <p>wrapped - phase wrapped into mod2<span class="math-tex">\(\pi\)</span> (<em>float32</em>)</p> </li> <li> <p>wrapcount - wrap count phase maps coded in the integer form (0,1,2...) (<em>uint8</em>)</p> </li> <li> <p>noised - wrapped phase images w/ synthetic noise (<em>float32</em>).</p> </li> </ul> </li> <li> <p><strong>Trained models</strong> - trained model files. These model files are in the format &quot;<strong>.h5</strong>&quot;, which contains the model architecture and the weights. They have been developed and saved with the <em>keras</em> library, and are loaded with the <em>keras.models.load_model()</em> function. The models list:</p> <ul> <li> <p><em>Unet_Denoising.h5 </em>- U-Net model used for denoising as an image translation task. The input is a wrapped phase image with noise and the output is the same wrapped phase distribution, but denoised. Model is trained on the synthetic phase dataset.</p> </li> <li> <p><em>Attn_Unet_Unwrapping.h5 </em>- U-Net model with Attention Gates and Residual Blocks trained for the semantic segmentation task. The input of the model is the wrapped phase image and its output is the wrap count map. Model is trained on the real phase dataset.</p> </li> </ul> </li> </ol> <p>&nbsp;</p> <p>&nbsp;</p> <p>[1] M. Baczewska, W. Krauze, A. Kuś, P. Stępień, K. Tokarska, K. Zukowski, E. Malinowska, Z. Brz&oacute;zka, and M. Kujawińska, &ldquo;On-chip holographic tomography for quantifying refractive index changes of cells&rsquo; dynamics,&rdquo; in Quantitative Phase Imaging VIII, vol. 11970 Y. Liu, G. Popescu, and Y. Park, eds., International Society for Optics and Photonics (SPIE, 2022), p. 1197008.<br> [2] P. Stępień, M. Ziemczonok, M. Kujawińska, M. Baczewska, L. Valenti, A. Cherubini, E. Casirati, and W. Krauze, &ldquo;Numerical refractive index correction for the stitching procedure in tomographic quantitative phase imaging,&rdquo; Biomed. Opt. Express 13, 5709&ndash;5720 (2022).<br> [3] M. Ziemczonok, A. Kuś, P. Wasylczyk, and M. Kujawińska, &ldquo;3d-printed biological cell phantom for testing 3d quantitative phase imaging systems,&rdquo; Sci. Reports 9, 1&ndash;9 (2019).<br> [4] M. Ziemczonok, A. Kuś, and M. Kujawińska, &ldquo;Optical diffraction tomography meets metrology &mdash; measurement accuracy on cellular and subcellular level,&rdquo; Measurement 195, 111106 (2022).<br> [5] W. Krauze, A. Kuś, M. Ziemczonok, M. Haimowitz, S. Chowdhury, and M. Kujawińska, &ldquo;3d scattering microphantom sample to assess quantitative accuracy in tomographic phase microscopy techniques,&rdquo; Sci. Reports 12, 1&ndash;9 (2022).</p>

opencc-by-4.0Jan 2023View details →
zenodo48/100

Data and Models from the study entitled, "Large-area automatic detection of shoreline stranded marine debris using deep learning"

<p>This repository contains data and models used in the study entitled, "Large-area automatic detection of shoreline stranded marine debris using deep learning". This study can be accessed as an open access publication at the following location: https://doi.org/10.1016/j.jag.2023.103515.</p> <p>The data set is comprised of 1,587 images (512 pixels x 512 pixels) which contains 10,703 individual bounding box labels of marine debris objects. The imagery was collected over the State of Hawai'i in 2015 at 2 centimeter resolution (ground spacing distance).</p> <p>The classification scheme consists of 8 labeled classes: unidentified object, processed wood, metal, vessel, net/cloth, buoy, tire, and line fragments.</p>

opencc-by-4.0Sep 2023View details →
zenodo48/100

MATLAB codes for : "Diagnosis and Prognosis of Faults in High-Speed Aeronautical Bearings with a Collaborative Selection Incremental Deep Transfer Learning Approach".

<p>The package contains all the materials needed to reproduce the findings of our paper. The paper is published by MDPI Applied Sciences journal and its details are as follow.</p> <p>Berghout, T.; Benbouzid, M. Diagnosis and Prognosis of Faults in High-Speed Aeronautical Bearings with a Collaborative Selection Incremental Deep Transfer Learning Approach.&nbsp;<em>Appl. Sci.</em>&nbsp;<strong>2023</strong>,&nbsp;<em>13</em>, 10916. https://doi.org/10.3390/app131910916</p> <p>1) Please you need to download the dataset from original link provided by introductory paper (Please read the above paper to find out about the datset used).<br> 2) Put the data in folders &quot;RawData&quot; for both experments.<br> 3) Please run the files for each experiment as provided, in alphabetical order.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Deep reinforcement learning for the control of microbial co-cultures in bioreactors

<p>Data for the figures in the paper:<br> <a href="https://www.biorxiv.org/content/10.1101/457366v2">https://www.biorxiv.org/content/10.1101/457366v2</a><br> (in press PLoS Comp Biol.)</p> <p>Abstract:<br> Multi-species microbial communities are widespread in natural ecosystems. When employed for biomanufacturing, engineered synthetic communities have shown increased productivity in comparison with monocultures and allow for the reduction of metabolic load by compartmentalising bioprocesses between multiple sub-populations. Despite these benefits, co-cultures are rarely used in practice because control over the constituent species of an assembled community has proven challenging. Here we demonstrate, in silico, the efficacy of an approach from artificial intelligence &ndash; reinforcement learning &ndash; for the control of co-cultures within continuous bioreactors. We confirm that feedback via reinforcement learning can be used to maintain populations at target levels, and that model-free performance with bang-bang control can outperform a traditional proportional integral controller with continuous control, when faced with infrequent sampling. Further, we demonstrate that a satisfactory control policy can be learned in one twenty-four hour experiment by running five bioreactors in parallel. Finally, we show that reinforcement learning can directly optimise the output of a co-culture bioprocess. Overall, reinforcement learning is a promising technique for the control of microbial communities.</p>

opencc-by-4.0Mar 2020View details →
zenodo44/100

Video classification using deep learning

<p>Material accompanying the paper &quot;Frame-by-frame annotation of video recordings using deep neural networks&quot;. Contains a selection of videos used in the paper, manual annotations, and results. Code is contained in the associated GitHub repository.&nbsp;See the readme for details.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Developing a deep Learning network to retrieve ocean hydrographic profiles in the North Atlantic from combined satellite and in situ measurements: test datasets.

<p>We provide here the datasets used for the test and assessment of a deep learning algorithm which is presently candidate for the development of a daily 3D ocean product covering the North Atlantic at 1/10&deg; resolution, over the 2010-2018 period, as part of the European Space Agency World Ocean Circulation project (ESA-WOC). The method is based on a stacked Long Short-Term Memory neural network, coupled to a Monte-Carlo dropout approach, and allows to project satellite-derived sea surface temperature, sea surface salinity and absolute dynamic topography data at depth after training with sparse co-located in situ vertical hydrographic profiles (Buongiorno Nardelli, 2020, doi:<a href="https://www.researchgate.net/deref/http%3A%2F%2Fdx.doi.org%2F10.3390%2Frs12193151?_sg%5B0%5D=0xE-347r7Hvb80klJcEo811AhUiXq-twG_E6l4yB-BfIKkVtW-lVLGcO02mTFkUczvozYYI0WCPyUBFR3kzWNGGZKg.ftvLheFrzHIJriO4qW2bdxalvR_TWt3MpwUfvto3EemhRgvDRGwJ9Mdy4Xr0IcGCfICivf4j-VqTgKxVvXRogA">10.3390/rs12193151</a>).&nbsp;</p> <p>The test dataset presented here includes different sets of co-located temperature and salinity vertical profiles:&nbsp;</p> <ul> <li>in situ observations extracted from the quality controlled Argo and CTD profiles produced by&nbsp;Copernicus Marine Environment Monitoring Service&nbsp;CORA 5.2 (<a href="http://marine.copernicus.eu/services-portfolio/access-to-products/">http://marine.copernicus.eu/services-portfolio/access-to-products/</a>,&nbsp;product_id: INSITU_GLO_TS_REP_OBSERVATIONS_013_001_b, doi: 10.17882/46219TS1,&nbsp;Szekely et al., 2019)&nbsp;and interpolated through a spline on a regularly spaced vertical grid (with 10 m intervals);</li> <li>climatological profiles extracted from World Ocean Atlas 2013 optimally interpolated monthly fields&nbsp;(Locarnini et al., 2013; Zweng et al., 2013), interpolated through a spline on a regularly spaced vertical grid (with 10 m intervals), upsized to a 1/10&deg; horizontal grid through a cubic spline and linearly interpolated in time between the central day of each month;</li> <li>synthetic profiles obtained through three different techniques: multivariate EOF reconstruction, a 2 layer feed-forward network (with 1000 units in each hidden layer) and a stacked LSTM network (with 2 LSTM layers and 35 hidden units)</li> </ul> <p><em>References:</em></p> <p>Buongiorno Nardelli, B.:&nbsp;A Deep Learning network to retrieve ocean hydrographic profiles from combined satellite and in situ measurements, 2020, <em>submitted</em>.</p> <p>Locarnini, R. A., Mishonov, A. V., Antonov, J. I., Boyer, T. P., Garcia, H. E., Baranova, O. K., Zweng, M. M., Paver, C. R., Reagan, J. R., Johnson, D. R., Hamilton, M. and Seidov, D.: World Ocean Atlas 2013. Vol. 1: Temperature., S. Levitus, Ed.; A. Mishonov, Tech. Ed.; NOAA Atlas NESDIS, 73(September), 40, doi:10.1182/blood-2011-06-357442, 2013.</p> <p>Szekely, T., Gourrion, J., Pouliquen, S. and Reverdin, G.: The CORA 5.2 dataset for global in situ temperature and salinity measurements: Data description and validation, Ocean Sci., 15(6), 1601&ndash;1614, doi:10.5194/os-15-1601-2019, 2019.</p> <p>Zweng, M. M., Reagan, J. R., Antonov, J. I., Mishonov, A. V., Boyer, T. P., Garcia, H. E., Baranova, O. K., Johnson, D. R., Seidov, D. and Bidlle, M. M.: World Ocean Atlas 2013, Volume 2: Salinity, NOAA Atlas NESDIS, 119(1), 227&ndash;237, doi:10.1182/blood-2011-06-357442, 2013.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record