Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,273

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,273 results for “Deep Learning”

Learn how ShareScore rates datasets ↗
zenodo44/100

DeepLabCut: markerless pose estimation of user-defined body parts with deep learning

<p>This data entry contains <strong>annotated mouse data from the <a href="https://www.nature.com/articles/s41593-018-0209-y">DeepLabCut Nature Neuroscience paper</a></strong>.</p> <p>This data entry contains a public release of annotated mouse data from the DeepLabCut paper. The trail-tracking behavior is part of an investigation into odor guided navigation, where one or multiple wildtype (C57BL/6J) mice are running on a paper spool and following&nbsp; odor&nbsp; trails. These experiments were carried out by Alexander Mathis &amp; Mackenzie Mathis in&nbsp; the Murthy lab at Harvard University. &nbsp;</p> <p>Data&nbsp; was&nbsp; recorded by&nbsp; two&nbsp; different&nbsp; cameras&nbsp; (640&times;480&nbsp; pixels with Point Grey Firefly (FMVU-03MTM-CS),&nbsp; and&nbsp; at&nbsp; approximately 1,700&times;1,200 pixels with Grasshopper 3 4.1MP Mono USB3 Vision (CMOSIS CMV4000-3E12)) at 30 Hz.&nbsp; The latter images were&nbsp; cropped around mice to generate images that are approximately 800&times;800. &nbsp;</p> <p>Here we share 1066, frames from multiple experimental sessions observing 7 different mice.&nbsp;&nbsp; Pranav Mamidanna labeled the snout, the tip of the left and right ear as well as the base of the tail in the example images. The data is organized in <a href="https://www.nature.com/articles/s41596-019-0176-0">DeepLabCut 2.0 project structure</a> with images and annotations in the labeled-data folder.&nbsp; The names are pseudocodes indicating mouse id and session id, e.g. m4s1 = mouse 4 session 1.</p> <p>Code for loading, visualizing &amp; training deep neural networks available at <a href="http://https://github.com/DeepLabCut/DeepLabCut"> https://github.com/DeepLabCut/DeepLabCut</a>.</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

A Deep Learning Dataset for Tomato Pest Leafminer TUTA ABSOLUTA

<p>The images of&nbsp;tomato leafminer (<em>Tuta absoluta</em>) were taken&nbsp;in in-house plots between August 2018 and May 2019 in&nbsp;Arusha, Tanzania.&nbsp;&nbsp;Under net-house that were controlled from other others. <em>T.absoluta</em> larvae were inoculated on the commonly grown&nbsp;tomato&nbsp;varieties at the early growth stage (herein, on the second day after transplanting). The images were taken for the first 2 weeks after inoculation. Images captured the canopy of the plants.&nbsp;</p> <p><strong>File Description</strong><br> All Images are in the <strong>.zip</strong> files; &quot;dataset_1_H.zip&quot;&nbsp;has 1926 Images, dataset_1_NH.zip has 325 Images, dataset_2 .zip has 3482 Images and the files labels are in &quot;file_labels.csv&quot; the image file name in column &quot;FileName&quot; and respective label in column &quot;Label&quot;, labels meaning&nbsp;&quot;1&quot; refer to healthy (plants not inoculated with <em>T.absoluta</em> larvae&nbsp;and &quot;2&quot; refer to <em>T.absoluta</em> affected plants.&nbsp; A total of 4341 image files are labelled.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Multi-fidelity Generative Deep Learning Turbulent Flows

<p>Data sets for the two numerical examples in the paper&nbsp;<a href="https://arxiv.org/abs/2006.04731">Multi-fidelity Generative Deep Learning Turbulent Flows</a>&nbsp;as well as two pre-trained models.&nbsp;&nbsp;In this work, a novel multi-fidelity deep generative model is introduced for the surrogate modeling of high-fidelity turbulent flow fields given the solution of a computationally inexpensive but inaccurate low-fidelity solver. The resulting surrogate is able to generate physically accurate turbulent realizations at a computational cost magnitudes lower than that of a high-fidelity simulation. The deep generative model developed is a conditional invertible neural network, built with normalizing flows, with recurrent LSTM connections that allow for stable training of transient systems with high predictive accuracy. Data is provided from OpenFOAM LES simulations for turbulent flow over backwards step and flow around an array of cylinders.</p> <p>Data-set Files:</p> <ul> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/backward_step_testing.tar.gz?versionId=320a523f-0015-4ba3-8c6e-66733ab5a1af">backward_step_testing.tar.gz</a>&nbsp;- Backward step testing data.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/backward_step_training.tar.gz?versionId=ac34ac15-973d-4fb8-8881-faa17eced69f">backward_step_training.tar.gz</a>&nbsp;- Backward step training data.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/cylinder_array_testing.tar.gz?versionId=ebbea725-c8b6-4338-978b-dc73f943552e">cylinder_array_testing.tar.gz</a>&nbsp;- Cylinder array testing data.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/cylinder_array_training.tar.gz?versionId=b0409c34-fb19-45c4-bbba-6968fdcfc4d8">cylinder_array_training.tar.gz</a>&nbsp;- Cylinder array training data.</li> </ul> <p>Pre-trained Models:</p> <ul> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/bstepWorkspace400.zip">bstepWorkspace400.zip</a>&nbsp;- Backward step&nbsp;pre-trained model.</li> <li><a href="https://zenodo.org/api/files/cf158661-2a30-4c8f-aa97-cd5c1600910c/cylinderWorkspace400.zip">cylinderWorkspace400.zip</a>&nbsp;- Cylinder array pre-trained model.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Galaxy Zoo DECaLS: Detailed Visual Morphology Measurements from Volunteers and Deep Learning for 314,000 Galaxies

<p>This repository contains the data released in the paper &quot;Galaxy Zoo DECaLS: Detailed Visual Morphology Measurements from Volunteers and Deep Learning for 314,000 Galaxies&quot; <em>(DOI to follow on publication).</em></p> <p>We release detailed morphology catalogues, both volunteer and automated, for Galaxy Zoo DECaLS.</p> <p>- gz_decals_volunteers_1_and_2 contains volunteer classifications for galaxies classified during the GZD-1 and GZD-2 campaigns.</p> <p>- gz_decals_volunteers_5 similarly contains classifications from the GZD-5 campaign. Note that GZD-5 used a modified schema designed to better detect mergers and weak bars, and includes many galaxies with only approx. five volunteer responses.</p> <p>- gz_decals_auto_posteriors contains the predicted posteriors for volunteer responses to all galaxies used in any campaign. The full posteriors are recorded as Dirichlet distribution concentrations. gz_decals_auto_posteriors also summarises these posteriors as the automated equivalent of previous Galaxy Zoo data releases;<strong> the expected vote fractions (mean posteriors)</strong>. Note that not all posteriors/vote fractions are relevant for every galaxy; we suggest assessing relevance using the estimated fraction of volunteers that would have been asked each question.</p> <p>We include a schema document, schema.md, to define the column names in each catalogue.</p> <p>We also release the galaxy images shown to volunteers on www.galaxyzoo.org during GZD-5. The images on which the automated classifier was trained may be derived from these volunteer-facing images. These images are split into four zip files, each of which contains images named by iauname inside a subfolder named by the first four characters in their iauname. Not all images were labelled during GZD-5 - refer to the catalog for training labels. We are working with the Zenodo team to add these large files to this repository - meanwhile, you can download them from The University of Manchester <a href="https://docs.google.com/document/d/1YgpnxiSJ7ffOW6FY8pX0pw93LTu8rLIdPL2PYhxW1fo/edit?usp=sharing">here</a>.</p> <p>The .csv and .parquet files contain identical data. Parquet is a fast column-oriented binary format which can be read with pd.read_parquet(loc, columns=[some columns]).</p> <p>You may also be interested in the <a href="https://github.com/mwalmsley/zoobot">github repository</a> which contains code to reproduce the model and to fine-tune it for new tasks (including pretrained weights).</p> <p>We will release updates if needed via Zenodo versioning. We recommend using the latest version of this repository. You can check the version you are currently viewing on the right-hand sidebar.</p> <p>Please cite the paper (DOI to follow on publication) when using the data in this repository.</p> <p>---</p> <p>History</p> <p>v0.0.1 (submission) provides the catalog files.</p> <p>v0.0.2 (first revision) renames the catalog files, adds flags for poorly sized galaxies, and includes the galaxy images via the University of Manchester</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Deep learning based automatic grounding line delineation in DInSAR interferograms

<p>This dataset contains a small subset of the AIS_cci GLL product, which covers several key glaciers and the corresponding HED-delineated grounding lines generated from our automatic delineation pipeline. A description of the attributes of the AIS_cci GLL product is provided in the&nbsp;<a href="https://climate.esa.int/media/documents/ST-UL-ESA-AISCCI-PUG-0001.pdf" target="_blank" rel="noopener">Product User Guide</a>.&nbsp; We do not indicate the split of the interferograms into training, validation and test sets as the complete AIS_cci dataset is not open-access.</p> <p>We also provide eight double difference interferograms at 100 m pixel size to demonstrate the generation of the features stack. Please note, eight samples are not sufficient to train the neural network to achieve the delineation capability described in our work.</p> <p>The "UUID" attribute in both GeoJSON files is an identifier that links the vector geometries to the interferogram TiFF files.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Deep learning to extract the meteorological by-catch of wildlife cameras: Supporting data, models and code

<p>This repository contains the data, models and code to train and deploy deep learning models related to the paper "Deep learning to extract the meteorological by-catch of wildlife cameras" published in the journal Global Change Biology (<a href="https://doi.org/10.1111/gcb.17078"><strong>https://doi.org/10.1111/gcb.17078</strong></a>).</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Deep learning segmentation projects of FIB-SEM dataset of U2-OS cell

<p>This submission includes ground truth datasets that were used to segment the nuclear envelope (NE), mitochondria, endoplasmic reticulum (ER) and Golgi from a human bone osteosarcoma epithelial cell (U2-OS) imaged using focused-ion beam scanning electron microscopy (FIB-SEM).</p><p>The full FIB-SEM dataset is deposited to EMPIAR (<a href="https://www.ebi.ac.uk/empiar">https://www.ebi.ac.uk/empiar</a>, EMPIAR-11746).&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Data for Paper "Scalable Semantic 3D Mapping of Coral Reefs with Deep Learning"

<p><strong>Example Data for DeepReefMap</strong></p> <p>This dataset contains input videos in MP4 format taken with GoPro Hero 10 Cameras in Reefs in the Red Sea to demonstrate the DeepReefMap tool, which is described in the paper "Scalable Semantic 3D Mapping of Coral Reefs with Deep Learning" by Sauder et al.</p> <p>It contains a directory for model checkpoints for semantic segmentation, and for the 3D SLAM component:</p> <p>```<br>checkpoints/<br>&nbsp; &nbsp; &nbsp; &nbsp; segmentation_net.pth<br>&nbsp; &nbsp; &nbsp; &nbsp; sfm_net.pth<br>```</p> <p>It also contains videos to run the reconstruction with. See the detailed instructions for running reconstructions in https://github.com/josauder/mee-deepreefmap</p> <p>```<br>input_videos/<br>&nbsp; &nbsp; &nbsp; &nbsp; GX_SINGLE_VIDEO.MP4<br>&nbsp; &nbsp; &nbsp; &nbsp; GX_VIDEO_1_OF_2.MP4<br>&nbsp; &nbsp; &nbsp; &nbsp; GX_VIDEO_2_OF_2.MP4<br>```</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

OCTDL: Optical Coherence Tomography Dataset for Image-Based Deep Learning Methods

<p>Optical coherence tomography (OCT) is a non-invasive imaging technique that has extensive clinical applications in ophthalmology. OCT enables the visualization of the retinal layers, playing a vital role in the early detection and monitoring of retinal diseases. OCT uses the principle of light wave interference to create detailed images of the retinal microstructures, making it a valuable tool for diagnosing ocular conditions. Optical Coherence Tomography Dataset for Image-Based Deep Learning Methods (OCTDL) comprising over 2000 OCT images labeled according to disease group and retinal pathology.</p> <p>The dataset consists of the following categories and images:<br>- Age-Related Macular Degeneration - 1231 images;<br>- Diabetic Macular Edema - 147 images;<br>- Epiretinal Membrane- 155 images;<br>- Normal - 332 images;<br>- Retinal Artery Occlusion - 22 images;<br>- Retinal Vein Occlusion - 101 images;<br>- Vitreomacular Interface Disease - 76 images.</p> <p>This dataset is published to provide researchers and developers with access to a large set of labeled images, which contributes to the development and improvement of algorithms for the automatic processing and analysis of OCT images for early diagnosis and monitoring of eye diseases. CSV file consists of file_name, disease, subcategory, condition, patient_id, eye, sex, year, image_width, and image_height. The dataset will be updated periodically.</p> <p>&nbsp;</p> <p>For more information and details about the dataset see:</p> <p>https://rdcu.be/dELrE</p> <p>https://arxiv.org/abs/2312.08255</p> <pre>@article{kulyabin2024octdl, title={OCTDL: Optical Coherence Tomography Dataset for Image-Based Deep Learning Methods}, author={Kulyabin, Mikhail and Zhdanov, Aleksei and Nikiforova, Anastasia and Stepichev, Andrey <br> and Kuznetsova, Anna and Ronkin, Mikhail and Borisov, Vasilii and Bogachev, Alexander <br> and Korotkich, Sergey and Constable, Paul A and Maier, Andreas}, journal={Scientific Data}, volume={11}, number={1}, pages={365}, year={2024}, publisher={Nature Publishing Group UK London},<br> doi={https://doi.org/10.1038/s41597-024-03182-7} } </pre>

opencc-by-4.0Dec 2023View details →
zenodo44/100

A Real-Time Eye-Tracking Dataset for Autism Severity Classification Using Deep Learning

<p>Eye-Tracking (ET) technologies have shown significant potential in autism research, providing critical insights into gaze patterns and their correlation with autism severity. However, a persistent challenge in developing Deep Learning (DL) models for ET analysis is the lack of publicly available, annotated datasets tailored for specific tasks. In order to close this gap, we present a novel, meticulously annotated resource designed to classify autism severity based on ET data. This dataset consists of 4,000 high-resolution (416&times;416 pixels) eye images derived from video recordings of 40 participants, evenly distributed across four autism severity groups: low, mild, medium, and high.</p> <p>Each participant's video was processed to extract 50 frames per session, capturing diverse gaze behaviors such as fixations, saccades, and smooth pursuits. Both left and right eye images were segmented from these frames, yielding 100 images per participant and ensuring balanced representation across severity categories (1,000 images per group). The dataset is annotated with detailed metadata, including subject ID, frame number, autism severity level, and eye type (left or right), providing a robust foundation for precise feature extraction and analysis.</p> <p><span>Facilitating its application in DL model development, this dataset addresses a critical gap in the limited availability of ET datasets. It provides a robust benchmark for autism severity classification, establishing a foundational resource for advancing Machine Learning(ML) research in the domain of autism</span><span>. This dataset serves as a critical resource for advancing ET-based classification models, fostering accurate and efficient assessment of autism severity, and supporting broader autism research.</span></p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Data for "Identification of 4876 Bent-Tail Radio Galaxies in the FIRST Survey using Deep Learning Combined with Visual Inspection"

<p>The data are the full versions of tables that will be published in the manuscript titled "Identification of 4876 Bent-Tail Radio Galaxies in the FIRST Survey using Deep Learning Combined with Visual Inspection" by The Astrophysical Journal Supplement Series.</p> <p>The table file named "FIRST_bt_table1.csv" is the full table for "A catalog of 4876 BTRGs identified from VLA FIRST survey". &nbsp;</p> <p>The table file named "FIRST_bt_table2.csv" is the full table for "Cluster details for BTRGs". &nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo44/100

Deposition of data for developing deep learning models to assess crack width and self-healing progress in concrete (krkCMd)

<p>This is a deposition of data for developing deep learning models to assess crack width and self-healing progress in concrete [1]. It relates to an experimental study on the autogenous self-healing of high-strength concrete [2]. Concrete specimens were prepared, matured, cracked, and exposed to self-healing. High-resolution scanning of the specimen surface and scale-invariant image processing were performed, multiple grid lines crossing cracks were established, and brightness degree profiles were extracted. Then, manual measurements of the crack widths were obtained by an operator.</p> <p>The dataset comprises 19,098 records of brightness profiles, reference crack width measurements, and benchmark measurements by deep learning and analytic models. The source images, which were stacked and marked with grid lines, are provided. The considerable number of brightness profiles coupled with manual reference measurements make the dataset well suited for developing an image-based deep learning models or analytic algorithms for assessing crack widths in concrete.</p> <p>The deposited data includes:</p> <ul> <li>krkCMd_table.csv: delimited, comma-separated text file containing a dataset of 19,098 crack brightness degree profiles, reference crack width measurements by operator, and benchmark measurements by a deep CNN metasensor and by an analytic edge detector.</li> <li>krkCMd_images.zip: archive containing source image files in folders by test series:&nbsp;<br>-&nbsp;&nbsp; stacked images of cracks in subsequent stages of self-healing (.tif files),<br>-&nbsp; &nbsp;zip archives assigned to image stacks and containing sets of ImageJ data files .roi,<br>- &nbsp; ImageJ .roi files specifying the locations of grid lines in the images.</li> <li>krkCMd_scripts.zip: archive containing custom scripts supporting image preprocessing and computing benchmark variables.</li> </ul> <p><span>For details please see the <a href="https://doi.org/10.1038/s41597-025-04485-z">data descriptor [1]</a>. When referring to the data in publications please cite [1].</span></p> <p>[1] Jakubowski, J., Tomczak, K. Dataset for developing deep learning models to assess crack width and self-healing progress in concrete.&nbsp;<em>Sci Data</em>&nbsp;<strong>12</strong>, 165 (2025). https://doi.org/10.1038/s41597-025-04485-z</p> <p>[2] Jakubowski, J. &amp; Tomczak, K. Deep learning metasensor for crack-width assessment and self-healing evaluation in concrete.&nbsp;<em>Constr. Build. Mater.</em>&nbsp;<strong>422</strong>, 135768 (2024). https://doi.org/10.1016/j.conbuildmat.2024.135768</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

RNA-Protein Interaction Prediction Using Network-Guided Deep Learning

<p>RNA-protein interactions are critical to various life processes, including fundamental translation and gene regulation. Identifying these interactions is vital for understanding the mechanisms underlying life processes. Then, ZHMolGraph is an advanced pipeline that integrates graph neural network sampling strategy and unsupervised large language models to enhance binding predictions for novel RNAs and proteins.</p> <div>&nbsp;</div>

openmit-licenseJul 2024View details →
zenodo44/100

Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network

<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Dataset from paper "Canopy palm cover across the Brazilian Amazon forests mapped with airborne LiDAR data and deep learning"

<p><strong>Data and code from the paper:</strong></p> <p>Dalagnol, R., Wagner, F. H., Emilio, T., Streher, A. S., Galv&atilde;o, L. S., Ometto, J. P. H. B., &amp; Arag&atilde;o, L. E. O. C. (2022). Canopy palm cover across the Brazilian Amazon forests mapped with airborne LiDAR data and deep learning. Remote Sensing in Ecology and Conservation, 1&ndash;14. https://doi.org/10.1002/rse2.264</p> <p><strong>Link:</strong>&nbsp;<a href="https://doi.org/10.1002/rse2.264">https://doi.org/10.1002/rse2.264</a></p> <p>&nbsp;</p> <p><strong>This repository contains:</strong></p> <p><strong>1) model_train.R:</strong> This is the code to run the U-Net model in R language.</p> <p><strong>2) input.rar:</strong> Dataset of lidar canopy height model (CHM) images and masks (labels) patches of canopy palms obtained from four sites in the Brazilian Amazon.&nbsp;The images/masks&nbsp;have 128 x 128 pixels, where each pixel represents 0.5 m in the terrain. The dataset contains 2,269 images and masks, with close to 7,000 palms manually labelled.</p> <p><strong>3) unet_weights_best.h5:</strong> These are the best weights for the U-Net architecture achieved in the paper.</p> <p><strong>4) palm_stats.RData:</strong> Data frame with the lat/lon coordinates and palm metrics extracted for the 610 lidar sites in the Brazilian Amazon. (i) n_total is the number of palms, (ii) n_ha is the density of palms per hectare, (iii) crown_ metrics are based on the area of palm segments (in square meters), (iv) cover_total is the total area occupied by palms in the forest canopy (in square meters), (v)&nbsp;cover_rel is the relative cover of palms in the forest canopy (in percentage), (vi) height_ metrics are based on the height of palm segments (in meters), (vii) palm_height_dif_mean is the mean difference between palm height and local canopy height, and (viii) palm_height_dif_pvalue&nbsp;is the p-value assessing the statistical difference between the palm and canopy heights where 0 means no difference and -1/+1 means a negative/positive difference.</p> <p>&nbsp;</p> <p>If you need anything else, please contact the corresponding author: Ricardo Dalagnol (ricds@hotmail.com).</p> <p>&nbsp;</p> <p><strong>If you use these data, please cite the paper:</strong></p> <p>Dalagnol, R., Wagner, F. H., Emilio, T., Streher, A. S., Galv&atilde;o, L. S., Ometto, J. P. H. B., &amp; Arag&atilde;o, L. E. O. C. (2022). Canopy palm cover across the Brazilian Amazon forests mapped with airborne LiDAR data and deep learning. Remote Sensing in Ecology and Conservation, 1&ndash;14. https://doi.org/10.1002/rse2.264</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Outputs of the Jupyter Notebook - Detecting floating objects using Deep Learning and Sentinel-2 imagery

<p>The dataset contains the outputs of the notebook &quot;Detecting floating objects using Deep Learning and Sentinel-2 imagery&quot;&nbsp;published in the ocean modelling section of The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Modelling codebase</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Marc Ru&szlig;wurm (author), EPFL-ECEO,&nbsp;<a href="https://github.com/MarcCoru">@marccoru</a></p> </li> </ul>

opencc-by-4.0Jan 2022View details →
zenodo44/100

VirHunter: a deep learning-based method for detection of novel RNA viruses in plant sequencing data

<p>This storage contains 2&nbsp;archives: toy datasets to test the training of the VirHunter and weights of the&nbsp; fully trained VirHunter models for 3 host species &nbsp;(peach, grapevine, sugar beet) and&nbsp;for fragment sizes 500 and 1000.&nbsp; .</p> <p>The toy dataset consists of 3 archived files: &#39;viruses.fasta&#39;, &#39;host.fasta&#39;, &#39;bacteria.fasta&#39;.</p> <p>&#39;viruses.fasta&#39; contains 10000 randomly selected plant viruses from the virus dataset described in the paper.</p> <p>&#39;host.fasta&#39; consists of peach chromosome 2.</p> <p>&#39;bacteria.fasta&#39; consists of 10 bacterial genomes selected randomly:&nbsp;GCF_000284415, GCF_000590555, GCF_001548055, GCF_002795265, GCF_003330825,&nbsp;GCF_003957805, GCF_005845345,&nbsp;GCF_009176625,&nbsp;GCF_010748935, GCF_014681765</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Automatic labelling of HeLa "Kyoto" cells using Deep Learning tools

<p><strong>Name</strong>: Automatic labelling&nbsp;of HeLa &ldquo;Kyoto&rdquo; cells using Deep Learning tools</p> <p><strong>Data type</strong>: Microscopy images from the dataset &ldquo;<strong>HeLa &ldquo;Kyoto&rdquo; cells&nbsp;under the scope</strong>&rdquo;, Brightfield (BF), Digital Phase Contrast (DPC, either &ldquo;raw&rdquo; or &ldquo;square-rooted&rdquo;), Tubulin and H2B fluorescent channel, paired with their corresponding nuclei or cell/cyto label images.</p> <p><strong>Labels images</strong>: Labels images were generated using the script <em>&ldquo;prepare_trainingDataset_cellpose.ijm</em>&rdquo;.</p> <p>Briefly, for 5 defined time-points (1,10,50,100,150), channels of interest were duplicated, resaved and :</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; nuclei label images were obtained using <a href="https://github.com/stardist/stardist">StarDist</a> on H2B channel</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; cell label images were obtained using <a href="https://github.com/MouseLand/cellpose">Cellpose</a> on Tubulin and H2B channels</p> <p>A quick visual inspection of the resulting label images concluded that they were satisfying enough, despite certainly not being perfect.</p> <p>Notes :</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; This labelling&nbsp;strategy:</p> <p>o&nbsp;&nbsp; will not produce 100% accurate labels, but they might be more reproducible than labels generated by humans and are (definitely) much faster to obtain.</p> <p>o&nbsp;&nbsp; is <strong>NOT a recommended way of generating labels images</strong>, but for educational purposes.</p> <p>-&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The fluorescent channels are part of the dataset to ease the process of review of the labels and are NOT used for training. We generated the labels from the fluorescent channels to later predict labels from the BF or DPC channels only. As such, the fluorescent channels should not be &ldquo;reused&rdquo; with our labels during training.</p> <p><strong>File format</strong>: .tif (16-bit)</p> <p><strong>Image size</strong>: 540x540 (Pixel size: 0.299 nm)</p> <p>&nbsp;</p> <p><strong><em>NOTE</em></strong>: This dataset uses&nbsp;the &ldquo;HeLa &ldquo;Kyoto&rdquo; cells&nbsp;under the scope&rdquo; &nbsp;dataset (<a href="https://doi.org/10.5281/zenodo.6139958">https://doi.org/10.5281/zenodo.6139958</a>) to automatically generate annotations</p> <p><strong><em>NOTE</em></strong>: This dataset was used to train cellpose models in the following Zenodo entry&nbsp;<a href="https://doi.org/10.5281/zenodo.6140111">https://doi.org/10.5281/zenodo.6140111</a></p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Data sets and models for the Deep API Learning Revisited paper

<p>Training and test data for the machine learning experiments described in the paper Deep API Learning Revisited paper.&nbsp; Trained models are also included.</p> <p>Deep API Learning Revisited paper:&nbsp;&nbsp;https://doi.org/10.1145/3524610.3527872</p> <p>GitHub repository:&nbsp;&nbsp;https://github.com/hapsby/deepAPIRevisited</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Hybrid Deep Learning Techniques for Securing Bioluminescent Interfaces in Internet of Bio Nano Things

<p>The data-set presents normal and anomalous values of twelve traffic parameters, generated by <strong>Bioluminescent bio-cyber Interfacing </strong>(BBI) in the I<strong>nternet of Bio Nano Things </strong>(IoBNT) based systems.</p> <p>The traffic parameters included in the data-set represent bio-electric and electro-bio transduction unit operation of BBI incorporating normal, as well as abnormal data to train and test machine/deep learning classifiers in discriminating attack scenarios.</p> <p>The parameters considered include the following: <strong>Cumulative concentration of released molecules, Elimination rate, Michaelis-Menten constant, Kinetic constant, Forward rate constant, Catalytic reaction constant, Ligand-receptor binding constant, Concentration of ATP, Concentration of information molecules, Release rate Reverse kinetic constant,</strong> and <strong>Reverse forward rate constant.</strong></p> <p>The data set is divided into training and testing data for simplified analysis, and application.</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record