Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

98

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

98 results for “CNN”

Learn how ShareScore rates datasets ↗
edi60/100

Evaluation of Mask R-CNN Model for Counting Reproductive Structures of Six Plant Species 1895-2018

Phenology––the timing of life-history events––is a key trait for understanding responses of organisms to climate. The digitization and online mobilization of herbarium specimens is rapidly advancing our understanding of plant phenological response to climate and climatic change. The current common practice of manually harvesting data from individual specimens greatly restricts our ability to scale data collection to entire collections. Recent investigations have demonstrated that machine-learning models can facilitate data collection from herbarium specimens. However, present attempts have focused largely on simplistic binary coding of reproductive phenology (e.g., flowering or not). Here, we use crowd-sourced phenological data of numbers of buds, flowers, and fruits of more than 3000 specimens of six common wildflower species of the eastern United States (Anemone canadensis, A. hepatica, A. quinquefolia, Trillium erectum, T. grandiflorum, and T. undulatum} to train a model using Mask R-CNN to segment and count phenological features. A single global model was able to automate the binary coding of reproductive stage with greater than 90% accuracy. Segmenting and counting features were also successful, but accuracy varied with phenological stage and taxon. Counting buds was significantly more accurate than flowers or fruits. Moreover, botanical experts provided more reliable data than either crowd-sourcers or our Mask R-CNN model, highlighting the importance of high-quality human training data. Finally, we also demonstrated the transferability of our model to automated phenophase detection and counting of the three Trillium species, which have large and conspicuously-shaped reproductive organs. These results highlight the promise of our two-phase crowd-sourcing and machine-learning pipeline to segment and count reproductive features of herbarium specimens, providing high-quality data with which to study responses of plants to ongoing climatic change.

openCC0Dec 2023View details →
zenodo52/100

Dataset for Accuracy of Grid-Connected Photovoltaic Power Plant: A Novel Approach Using Hybrid Variational Mode Decomposition and CNN-LSTM Model

<p>This research paper introduces a deep learning hybrid model employing Convolutional Neural Network Long Short-Term Memory (CNN-LSTM) for short-term photovoltaic (PV) solar energy forecasting.The proposed method integrates the Variational Mode Decomposition (VMD) algo-rithm with the CNN-LSTM model to predict PV power generation from a solar farm in Boussada, Algeria, from January 1, 2019, to December 31, 2020. The performance of the developed model is benchmarked against other deep learning models (VMD-CNN, VMD-LSTM, CNN-LSTM) across various time horizons (15, 30, and 60 minutes) to provide a comprehensive evaluation. Our findings exhibit greater performance of the developed model compared to other architectures, showcasing promising results in solar power forecasting. This research contributes to the main goal of enhancing EMS by providing accurate solar energy forecasts.</p>

opencc-by-4.0Mar 2024View details →
zenodo48/100

Leaf Vein Network CNN Results

<p>Results for leaf vein networks extracted using the LeafVeinCNN software package. The original image data set is available from Blonder et al. (2019)&nbsp;<a href="https://doi.org/10.1002/ecy.2844">https://doi.org/10.1002/ecy.2844</a>. The LeafVeinCNN software used in the analysis is available at&nbsp;<a href="https://doi.org/10.5281/zenodo.4007731"> </a><a href="https://doi.org/10.5281/zenodo.4007730">https://doi.org/10.5281/zenodo.4007730</a></p> <ul> <li>The Results_xxx.zip files contain&nbsp;all the Excel results spreadsheets separated by the code for each field site.</li> <li>results.xls provides a summary of all the network metrics for each file that was analysable</li> <li>Results_figures.pdf provides a summary image of the processing steps and results for each leaf segment</li> <li>Network_images.pdf contains a colour-coded image of each network superimposed on the leaf segment</li> <li>HLD_plots shows the binary tree following Hierarchical Network Decomposition</li> <li>PR_results.zip contains the Excel spreadsheets for evaluation of different enhancement methods for each leaf segment.</li> <li>PR_summary.xls provides a summary of the performance of each enhancement method.</li> <li>PR_F1_images.pdf and PR_FBeta2_images.pdf show the pixel classification for each enhancement and segmentation method&nbsp;compared to the manual ground-truth using two different optimum criteria (F1 and FBeta2).</li> <li>PR_fullwidth_plots show the full Precision-Recall plots for the full-width binary image compared to the manual ground-truth using the FBeta2 metric.</li> <li>PR_skeleton_plots show the full Precision-Recall plots for the skeletonised binary image&nbsp;compared to the manual ground-truth&nbsp;using the FBeta2 metric.</li> <li>PR_threshold_plots.pdf show how a set of network metrics vary with the segmentation threshold for each enhancement method.</li> </ul>

opencc-by-4.0Aug 2020View details →
zenodo48/100

Hail Event on 2022-06-28 in Locarno-Monti (TI), Switzerland: Drone Photogrammetry Imagery, Mask R-CNN Model and Analysis Data of Hailstones

<p>This hail data collection belongs to a drone hail survey performed on 2022-06-28 in Locarno-Monti (TI, Switzerland). The supercell reached the location around 07:50 UTC in the morning. Only one photogrammetry flight could be performed and thus no estimation of the hail melting process is available. The orthophoto is masked to ignore parts where detection of hail is unwanted.</p> <p>&nbsp;</p> <p>Expert 1 (lai, mlainer), Expert 2 (jtm), Expert 3 (por, jportmann)</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Transfer learning for galaxy feature detection: Finding Giant Star-forming Clumps in low redshift galaxies using Faster R-CNN

<p>This repository contains the data released in the paper 'Transfer learning for galaxy feature detection: Finding Giant Star-forming Clumps in low redshift galaxies using Faster R-CNN'&nbsp;<em>(DOI: <a href="https://doi.org/10.1093/rasti/rzae013">10.1093/rasti/rzae013</a>).</em></p> <p>We release a detailed catalogue of Giant Star-forming Clumps (GSFCs), detected for the full set of Galaxy Zoo: Clump Scout&nbsp;galaxies observed by SDSS using the Faster R-CNN architecture with the Zoobot classification-CNN as a feature extraction backbone.</p> <p>The final models and code are made publicly available via Github:&nbsp;<a href="https://github.com/ou-astrophysics/Faster-R-CNN-for-Galaxy-Zoo-Clump-Scout">https://github.com/ou-astrophysics/Faster-R-CNN-for-Galaxy-Zoo-Clump-Scout</a>.</p> <p>We will release updates if needed via Zenodo versioning. We recommend using the latest version of this repository. You can check the version you are currently viewing on the right-hand sidebar.</p> <p>Please cite the paper (DOI: <a href="https://doi.org/10.1093/rasti/rzae013">10.1093/rasti/rzae013</a>) when using the data in this repository.</p> <p>The csv-file <em>FRCNN_Zoobot_SDSS_GZCS_detections.csv</em>&nbsp;has the following columns. Alternatively, the file <em>FRCNN_Zoobot_SDSS_GZCS_detections.gzip</em> contains the same data but stored as a parquet-file.</p> <table> <tbody><tr> <th>Column name</th> <th>Description</th> </tr> </tbody><tbody> <tr> <td>specobjid</td> <td>SDSS spec object ID</td> </tr> <tr> <td>dr7objid</td> <td>SDSS DR7 object ID</td> </tr> <tr> <td>clump_id</td> <td>Clump index</td> </tr> <tr> <td>clump_label_id</td> <td>Clump label ID (1 or 2)</td> </tr> <tr> <td>clump_label_name</td> <td>Clump label name</td> </tr> <tr> <td>clump_score</td> <td>Detection score for the clump</td> </tr> <tr> <td>clump_centre_ra</td> <td>Clump centroid RA in degrees</td> </tr> <tr> <td>clump_centre_dec</td> <td>Clump centroid dec in degrees</td> </tr> <tr> <td>clump_flux_u</td> <td>Clump u-band flux in Jy</td> </tr> <tr> <td>clump_flux_g</td> <td>Clump g-band flux in Jy</td> </tr> <tr> <td>clump_flux_r</td> <td>Clump r-band flux in Jy</td> </tr> <tr> <td>clump_flux_i</td> <td>Clump i-band flux in Jy</td> </tr> <tr> <td>clump_flux_z</td> <td>Clump z-band flux in Jy</td> </tr> <tr> <td>clump_flux_err_u</td> <td>Clump u-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_g</td> <td>Clump g-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_r</td> <td>Clump r-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_i</td> <td>Clump i-band flux error in Jy</td> </tr> <tr> <td>clump_flux_err_z</td> <td>Clump z-band flux error in Jy</td> </tr> <tr> <td>clump_mag_u</td> <td>Clump u-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_g</td> <td>Clump g-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_r</td> <td>Clump r-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_i</td> <td>Clump i-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_z</td> <td>Clump z-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_ext_mag_u</td> <td>Clump u-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_g</td> <td>Clump g-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_r</td> <td>Clump r-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_i</td> <td>Clump i-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_ext_mag_z</td> <td>Clump z-band extinction (E(B-V), AB-mag)</td> </tr> <tr> <td>clump_mag_corr_u</td> <td>Clump corrected u-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_g</td> <td>Clump corrected g-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_r</td> <td>Clump corrected r-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_i</td> <td>Clump corrected i-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_z</td> <td>Clump corrected z-band magnitude (AB-mag)</td> </tr> <tr> <td>clump_mag_corr_u_g</td> <td>Clump colour (u-g)</td> </tr> <tr> <td>clump_mag_corr_g_r</td> <td>Clump colour (g-r)</td> </tr> <tr> <td>clump_mag_corr_r_i</td> <td>Clump colour (r-i)</td> </tr> <tr> <td>clump_mag_corr_i_z</td> <td>Clump colour (i-z)</td> </tr> <tr> <td>clump_flux_ratio</td> <td>Est. clump/galaxy near-UV flux ratio (u-band)</td> </tr> <tr> <td>is_clump_3pct</td> <td>Flag (True/False) if clump/galaxy flux ratio is &gt;3%</td> </tr> <tr> <td>is_clump_8pct</td> <td>Flag (True/False) if clump/galaxy flux ratio is &gt;8%</td> </tr> <tr> <td>galaxy_ra</td> <td>Host galaxy RA in degrees</td> </tr> <tr> <td>galaxy_dec</td> <td>Host galaxy dec in degrees</td> </tr> <tr> <td>galaxy_z</td> <td>Host galaxy redshift</td> </tr> <tr> <td>galaxy_mag_u</td> <td>Host galaxy u-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_g</td> <td>Host galaxy g-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_r</td> <td>Host galaxy r-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_i</td> <td>Host galaxy i-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_z</td> <td>Host galaxy z-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_u</td> <td>Host galaxy u-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_g</td> <td>Host galaxy g-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_r</td> <td>Host galaxy r-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_i</td> <td>Host galaxy i-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_mag_err_z</td> <td>Host galaxy z-band magnitude error (AB-mag)</td> </tr> <tr> <td>galaxy_flux_u</td> <td>Host galaxy u-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_g</td> <td>Host galaxy g-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_r</td> <td>Host galaxy r-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_i</td> <td>Host galaxy i-band flux in Jy</td> </tr> <tr> <td>galaxy_flux_z</td> <td>Host galaxy z-band flux in Jy</td> </tr> <tr> <td>galaxy_expAB_r</td> <td>Host galaxy axis ratio from SDSS</td> </tr> <tr> <td>galaxy_expRad_r</td> <td>Host galaxy exponential fit scale radius from SDSS</td> </tr> <tr> <td>galaxy_lmass</td> <td>Host galaxy log mass in MSun</td> </tr> <tr> <td>galaxy_lssfr</td> <td>Host galaxy log specific SFR</td> </tr> <tr> <td>galaxy_mag_corr_u</td> <td>Host galaxy corrected u-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_g</td> <td>Host galaxy corrected g-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_r</td> <td>Host galaxy corrected r-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_i</td> <td>Host galaxy corrected i-band magnitude (AB-mag)</td> </tr> <tr> <td>galaxy_mag_corr_z</td> <td>Host galaxy corrected z-band magnitude (AB-mag)</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Fluorescently-labelled zebrafish pronephroi + ground truth classes (normal/cystic) + trained CNN model

<p>This upload contains :</p> <p>- <strong>images.zip:&nbsp;</strong> microscope images of fluorescently-labelled pronephroi in larvae of the <em>Tg(wt1b:EGFP)</em> transgenic zebrafish line showing 2 morphologies (normal vs cystic) upon injection with Co-Mo or ift172-MO, respectively. Images were obtained using &nbsp;an ACQUIFER Imaging Machine widefield high content screening microscope.</p> <p>Reference:&nbsp;</p> <p>Pandey, G., Westhoff, J., Schaefer, F. and Gehrig, J. (2019). <strong>A Smart Imaging Workflow for Organ-Specific Screening in a Cystic Kidney Zebrafish Disease Model</strong>. International Journal of Molecular Sciences <em>20</em>, 1290, doi:<a href="https://doi.org/10.3390/ijms20061290">10.3390/ijms20061290</a>.</p> <p>&nbsp;</p> <p>- <strong>Annotations-***.csv : </strong>Tables containing ground-truth category classes (normal vs cystic) for the images in the zip file.</p> <p>The tables contain&nbsp;columns with the image filename, folder and category.</p> <p>Note&nbsp;:&nbsp;<strong>the Folder column should be updated with the root folder directory once downloaded on your machine.</strong></p> <p>These&nbsp;files were&nbsp;generated with the Fiji plugin <em>single-class (button)</em>&nbsp;from the <em>Qualitative-Annotations</em> update site.</p> <p>The 2 files contain&nbsp;the same information, they only differ in the formatting&nbsp;of the category, the <em>singleColumn </em>file has a single category column while the <em>multiColumn</em> has 2 columns (normal/cystic) with 0/1 encoding.</p> <p>The choice of category encoding solely depends on how the table is used, i.e. in which training workflow, home-made script or software.</p> <p>- <strong>trainedModel.zip :&nbsp;</strong>This archive contains 2 files:&nbsp;<strong>(1) </strong>a h5 file corresponding to a trained deep-learning model to classify the images of the dataset in the 2 categories (normal vs cystic), and <strong>(2)</strong>&nbsp;a text file containing the class names. Both files are necessary to predict the category of new images similar to the one in the dataset, for instance using the published KNIME workflows.</p>

opencc-by-4.0Aug 2020View details →
zenodo44/100

Ortoimages for CNN trainning

<p>Preprocessed ortoimages from Castilla and Leon to use in the training of a neural network to find ruins of ancient roman camps among the fields.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

CNN-Filter-DB

<p><strong>A diverse database of over 1.4B 3x3&nbsp;convolution filters extracted from CNN models trained for various tasks in diverse image domains.</strong></p> <p>We collected a total of 647 publicly available CNN models&nbsp;that have been pre-trained for various 2D visual tasks. In order to provide a heterogeneous and diverse representation of convolution filters &quot;in the wild&quot;, we retrieved pre-trained models for 11 different tasks e.g. such as classification, segmentation} and image generation. We also recorded various meta-data such as depth and frequency of included operations for each model, and manually categorized the variety of used training sets into 16 visually distinctive groups like natural scenes, medical ct, seismic, or astronomy. In total, the models were trained on 71 different data sets. The dominant subset is formed by image classification&nbsp;models trained on ImageNet1k (355 models).</p> <p><strong>More details:</strong>&nbsp;<a href="https://github.com/paulgavrikov/cnn-filter-db">https://github.com/paulgavrikov/cnn-filter-db</a></p>

opencc-by-sa-4.0Jun 2022View details →
zenodo44/100

CNN-Filter-DB-Robust

<p>Dataset for the Paper &quot;Adversarial Robustness through the Lens of Convolutional Filters&quot;.</p> <p><strong>More details:</strong>&nbsp;<a href="https://github.com/paulgavrikov/cnn-filter-db">https://github.com/paulgavrikov/cvpr22w_RobustnessThroughTheLens</a></p>

opencc-by-sa-4.0Apr 2022View details →
zenodo44/100

CNN Wild Park - Graph Neural Networks for Learning Equivariant Representations of Neural Networks

<p>This repository contains the <strong>CNN Wild Park</strong> dataset from the paper:</p> <blockquote> <p><strong>Graph Neural Networks for Learning Equivariant Representations of Neural Networks</strong><br><a href="https://mkofinas.github.io/">Miltiadis Kofinas</a>*,&nbsp;<a href="https://bknyaz.github.io/">Boris Knyazev</a>, <a href="https://www.cyanogenoid.com/">Yan Zhang</a>,&nbsp;<a href="https://yunlu-chen.github.io/">Yunlu Chen</a>,&nbsp;<a href="https://gertjanburghouts.github.io/">Gertjan J. Burghouts</a>,&nbsp;<a href="https://egavves.com/">Efstratios Gavves</a>,&nbsp;<a href="https://www.ceessnoek.info/">Cees G. M. Snoek</a>,&nbsp;<a href="https://davzha.netlify.app/">David W. Zhang</a>*<br><em>ICLR 2024</em> (oral)<br><a href="https://arxiv.org/abs/2403.12143">https://arxiv.org/abs/2403.12143</a><br><a href="https://github.com/mkofinas/neural-graphs">https://github.com/mkofinas/neural-graphs</a><br>*Joint first and last authors</p> </blockquote> <p>We introduce a new dataset of CNNs, which we term <em>CNN Wild Park</em>.<br>The dataset consists of 117,241 checkpoints from 2,800 CNNs, trained for up to 1,000 epochs on CIFAR10.<br>The CNNs vary in the number of layers, kernel sizes, activation functions, and residual connections between arbitrary layers.</p> <p>More specifically, we construct the CNN Wild Park dataset by training 2,800 small CNNs with different architectures for 200 to 1,000 epochs on CIFAR10. We retain a checkpoint of its parameters every 10 steps and also record the test accuracy. The CNNs vary by:</p> <ul> <li>Number of layers L in [2, 3, 4, 5] (note that this does not count the input layer).</li> <li>Number of channels per layer c_l in [4, 8, 16, 32].</li> <li>Kernel size of each convolution k_l in [3, 5, 7].</li> <li>Activation functions at each layer are one of ReLU, GeLU, tanh, sigmoid, leaky ReLU, or the identity function.</li> <li>Skip connections between two layers with at least one layer in between. Each layer can have at most one incoming skip connection. We allow for skip connections even in the case when the number of channels differ, to increase the variety of architectures and ensure independence between different architectural choices. We enable this by adding the skip connection only to the min(c_n, c_m) nodes.</li> </ul> <p>We divide the dataset into train/val/test splits such that checkpoints from the same run are <strong>not</strong> contained in both the train and test splits.&nbsp;</p> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0May 2024View details →
zenodo44/100

Dataset: Mask R-CNN Based C. Elegans Detection with a DIY Microscope

<p>The dataset consists of images of C. elegans in Petri Dish that were&nbsp;captured at a frequency of 1 Hz at 3280 &times; 2464 pixels via a&nbsp; Raspberry Pi based DIY Microscope. Further details of the recording setup and the dataset can be found in the corresponding article.</p> <p>Up on use, please cite the following article&nbsp;<a href="https://doi.org/10.3390/bios11080257">https://doi.org/10.3390/bios11080257</a>&nbsp;such as:</p> <p>Fudickar, S.; Nustede, E.J.; Dreyer, E.; Bornhorst, J. Mask R-CNN Based C. Elegans Detection with a DIY Microscope.&nbsp;<em>Biosensors</em>&nbsp;<strong>2021</strong>,&nbsp;<em>11</em>, 257. https://doi.org/10.3390/bios11080257</p> <p>&nbsp;</p> <p><br> &nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Supporting Data -- Evaluating Mask R-CNN Models to Extract Terracing across Oceanic High Islands: an example from Sāmoa.

<p>This dataset provides supplemental information for the manuscript, &quot;Diverse terracing practices revealed by automated lidar analysis across the Sāmoan islands&quot;, submitted to Archaeological Prospection. The dataset&nbsp;contains a trained Mask R-CNN deep learning model designed for detecting archaeological terracing features on the islands of American Samoa, associated training data, and the raw and cleaned output of detected terraces.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

CNN YouTube: Titles, Views & Posting Dates Dataset

<p>This dataset was scrapped using one of the python modules called &quot;SiteScraper&quot;. The &quot;SiteScraper&quot; module is build on top of selenium. Here is the link to the module :&nbsp;<code>https://github.com/ibrahim-string/SiteScraper</code></p> <p>Text summarisation can be done on titles and correlation between views and titles can be explored using this dataset.</p> <p>This dataset will be very usefull To find out what kind of content democrats watch.</p> <p>This dataset will be updated every week.</p>

opencc-byJun 2023View details →
zenodo40/100

Convolutional Neural Net (CNN) models for ENCODE-Roadmap DNase-seq peaks and Transcription Factor ChIP-seq peaks - Basset architecture

<p>Deep learning models trained on epigenomic landscapes from ENCODE and Roadmap Epigenomics. The models are Basset convolutional neural networks (Kelley, et al 2016). The dataset used to train these models can be found at https://doi.org/10.5281/zenodo.4059038. The file `nn.encode-roadmap.models.basset.clf.tar.gz` contains 10 cross-validated models in Tensorflow framework files as well as details on the architecture, cross-validation scheme, and training of these models. The file `nn.encode-roadmap.models.basset.clf.np_weights.tar.gz` contains the 10 cross-validated models&#39; weights extracted to numpy array files (.npz).</p>

openmit-licenseSep 2020View details →
zenodo40/100

Convolutional Neural Net (CNN) models for epigenomic landscapes in epidermal differentiation - Basset architecture, classification and regression

<p>Deep learning models trained on epigenomic landscapes in keratinocyte differentiation. The models are Basset convolutional neural networks (Kelley, et al 2016). The dataset used to train these models can be found at https://doi.org/10.5281/zenodo.4062509. The file `nn.ggr.models.basset.clf.tar.gz` contains 10 cross-validated models that were pretrained using ENCODE-Roadmap trained model weights as initialization weights and also 10 cross-validated models that were initialized with random weights. Similarly, the file `nn.ggr.models.basset.regr.tar.gz` contains 10 cross-validated models that were pretrained using the classification model weights as initialization weights and also 10 cross-validated models that were initialized with random weights.</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

VGQ-CNN: Moving Beyond Fixed Cameras and Top-Grasps for Grasp Quality Prediction

<p>This dataset includes all the data and trained models to replicate our work for VGQ-CNN (accepted for IJCNN 2022). You can find the code to use this dataset on <a href="https://github.com/AuCoRoboticsMU/vgq-cnn">github</a>. To replicate the work done for VGQ-CNN, use the data in vgq-dset.zip. Trained models of VGQ-CNN, Fast-VGQ-CNN and GQ-CNN are available in VGQ-CNN_models.zip.</p> <p>&nbsp;</p> <p>To create your own, subsampled training and testing data, adjust our code on github to your subsampling constraints and use full_rendered_dset (created by unpacking full_rendered_dset_tensors.zip and full_rendered_dset_images.zip into the unpacked directory of full_rendered_dset_info.zip).</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

YOGData: Labelled data (YOLO and Mask R-CNN) for yogurt cup identification within production lines

<p><strong>D</strong><strong>ata abstract:</strong><br> The&nbsp;YogDATA dataset contains&nbsp;images from an industrial laboratory production line when it is functioned to quality yogurts.&nbsp;The case-study for the recognition of yogurt cups requires training of Mask R-CNN and YOLO v5.0 models with a set of corresponding images. Thus, it is important to collect the corresponding images to train and evaluate the class. Specifically, the&nbsp;YogDATA&nbsp;dataset includes the same labeled data for&nbsp;Mask R-CNN&nbsp;(coco format)&nbsp;and YOLO models. For the YOLO architecture, training&nbsp;and validation datsets&nbsp;include sets of images in jpg format&nbsp;and their annotations in txt file format. For the Mask R-CNN architecture, the annotation of the same sets of images are included in json file format&nbsp;(80% of images and annotations of each subset&nbsp;are in training set and&nbsp;20% of images of each subset are in test set.)&nbsp;<br> &nbsp;</p> <p><strong>Paper abstract:</strong><br> The explosion of the digitisation of the traditional industrial processes and procedures is consolidating a positive impact on modern society by offering a critical contribution to its economic development. In particular, the dairy sector consists of various processes, which are very demanding and thorough. It is crucial to leverage modern automation tools and through-engineering solutions to increase their efficiency and continuously meet challenging standards. Towards this end, in this work, an intelligent algorithm based on machine vision and artificial intelligence, which identifies dairy products within production lines, is presented. Furthermore, in order to train and validate the model,&nbsp;&nbsp;the YogDATA dataset was created that includes yogurt cups within a production line. Specifically, we evaluate two deep learning models (Mask R-CNN and YOLO v5.0) to recognise and detect each yogurt cup in a production line, in order to automate the packaging processes of the products. According to our results, the performance precision of the two models is similar, estimating its at 99\%.&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

CNN for the classification of ICE-CAMERA images of Antarctic ice particles

<p>-The file &#39;ZENODO_FILES.rar&#39; contains the GoogleNet Convolutional Neural Network (CNN) trained to classify pre-processed ICE-CAMERA images (224*224*3) into 14 classes. CNN was developed for (Mathworks) MATLAB&reg; R2020b.</p> <p>-The ICE-CAMERA images used for training, validation and testing the CNN are also contained in specific folders.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Single pixel s*t Landsat time series training data for CNN

<p>Single pixel s*t Landsat time series classification using 1D CNN</p> <p>Sep 22, 2022 update (version 2):&nbsp;<br> The 1D CNN classification codes are available at https://github.com/hankui/cnn_Landsat_time_series_classification_v2-Python</p> <p>The NLCD training data is available at 10.5281/zenodo.7106054</p> <p>The NLCD training data is derived from Landsat 5/7 analysis ready data (ARD) in year 2011 (as x predictor variable) and National Land Cover Database (NLCD) 2011 (as y response variable)</p> <p>The NLCD training data is distributed across Continental United States (CONUS) with 3,314,439 30m pixel locations</p> <p>The NLCD training data include (i) NLCD label with 15 classes, i.e., all NLCD classes except ice (https://www.mrlc.gov/data/legends/national-land-cover-database-class-legend-and-description)<br> &nbsp;&nbsp; &nbsp;(ii) year 2011 growing season Landsat ARD percentiles for Landsat 5/7 bands 2, 3, 4, 5 and 7 and for 8 band ratios derived from the five bands&nbsp;<br> &nbsp;&nbsp; &nbsp;(iii) percentiles include 10th, 20th, 25th, 30th, 35th, 40th, 50th (median), 60th, 65th, 70th, 75th, 80th, 90th so that&nbsp;<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;one can use 5 percentiles (10th, 25th, 50th, 75th, and 90th)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;7 percentiles (10th, 20th, 35th, 50th, 65th, 80th, and 90th)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;9 percentiles (10th, 20th, 30th, 40th, 50th, 60th, 70th, 80th, and 90th)<br> &nbsp;&nbsp; &nbsp;(iv) the pixel location represented in Landsat ARD tile h and v no. and the pixel i and j locations in the tile<br> &nbsp;&nbsp; &nbsp;(v) the no. of the cloud free observations in 2011 growing season derived for the pixel location<br> &nbsp;&nbsp;</p> <p>#*************************************************************************************************************#</p> <p>A munuscript describing how the data were derived and how the 1D CNN was adapted to the data is in review&nbsp;</p> <p><br> #*************************************************************************************************************#</p> <p>The codes were written in python (v3.7) and tensorflow (v2.6).&nbsp;</p> <p>The parameters are:</p> <p>(1) learning rate: cnn training initial learning rate 0.01 used in the paper&nbsp;</p> <p>(2) epoch: cnn training epochs 70 used in the paper&nbsp;</p> <p>(3) method: cnn training optimizer method 1: Adam method 2: dynamic learning rate used in the paper</p> <p>(4) L2: L2 regularization value; 0.001 used in paper&nbsp;</p> <p>(5) layer: no. of CNN layers (can be 4, 5 and 8) and 5 and 8 used in the paper</p> <p>(6) perc: training data percentages (can be 0.1, 0.5 and 0.9) tested in the paper; the evaluation is used the left 10%&nbsp;</p> <p>(7) gpui: which gpu process it will use (only applicable with multi-gpus)&nbsp;</p> <p>(8) IMG_HEIGHT: the no. of percentiles (can be 3, 5, 7 and 9) and 5, 7 and 9 used in the paper&nbsp;</p> <p>An example would be:&nbsp;</p> <p>version=7_4&nbsp;</p> <p>layer=5; perc=0.1; gpui=0;IMG_HEIGHT=5</p> <p>method=0; learning_rate=0.01; &nbsp; epoch=10; iter=1; L2=0.001; sleep ${SLEEP}; ## Hank layer=5; perc=0.1;&nbsp;</p> <p>echo &quot;python Pro_2d1d_CNN_v${version}.py ${learning_rate} ${epoch} ${method} ${L2} ${layer} ${perc} ${gpui} ${IMG_HEIGHT} &quot;</p> <p>python Pro_2d1d_CNN_v${version}.py ${learning_rate} ${epoch} ${method} ${L2} ${layer} ${perc} ${gpui} ${IMG_HEIGHT} &gt; layer${layer}.p${perc}.d${IMG_HEIGHT}.rate${learning_rate}.e${epoch}.L${L2}.v${version} &amp;&nbsp;</p> <p><br> #*************************************************************************************************************#</p> <p>Aug 29, 2021 (version 1):&nbsp;<br> Training data: There are 2 input text files (csv) storing the 3,314,439 NLCD and 484,476 CDL land cover&nbsp;training samples:<br> &nbsp;&nbsp;&nbsp;&nbsp;NLCD training: ./NLCD/metric.ard.nlcd.Mar01.18.40.txt<br> &nbsp;&nbsp;&nbsp;&nbsp;CDL training: ./CDL/metric.ard.nlcd.Mar01.18.40.txt</p> <p>The codes and their usages are at:&nbsp;<br> &nbsp;&nbsp; &nbsp;https://github.com/hankui/cnn_Landsat_time_series_classification_v1-R<br> &nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

CNN for Modeling Sanskrit Originated Bengali and Hindi Language Dataset

<p>Though recent works have focused on modeling high resource languages, the area is still unexplored for low resource languages like Bengali and Hindi. We propose an end-to-end trainable memory efficient CNN architecture named CoCNN to handle specific characteristics such as high inflection, morphological richness, flexible word order and phonetical spelling errors of Bengali and Hindi. In particular, we introduce two learnable convolutional sub-models at word and at sentence level that are end-to-end trainable. We show that state-of-the-art (SOTA) Transformer models including pretrained BERT do not necessarily yield the best performance for Bengali and Hindi. CoCNN outperforms pretrained BERT with 16X less parameters and achieves much better performance than SOTA LSTMs on multiple real-world datasets. This is the first study on the effectiveness of different architectures from Convolution, Recurrent, and Transformer neural net paradigm for modeling Bengali and Hindi.</p>

openmit-licenseOct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record