Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
129
datasets available to search
ShareScore release 0.7.1
Dataset results
129 results for “Deep Neural Networks”
Reconstructing Faces from fMRI Patterns using Deep Generative Neural Networks.
Open the record for dataset details and reuse information.
HEroBM: a deep equivariant graph neural network for high-fidelity backmapping from coarse-grained to all-atom structures
<p><span>Molecular simulations play a pivotal role in chemistry, biology, and material sciences, enabling the</span><br><span>study of complex dynamic properties within systems. Coarse-grained (CG) techniques have emerged</span><br><span>as indispensable tools in this domain, facilitating the sampling of large-scale systems and extending</span><br><span>simulation timescales by simplifying system representation. However, CG approaches involve a trade-</span><br><span>off: they sacrifice atomistic details that may be crucial for understanding the underlying processes.</span><br><span>To address this challenge, a recommended strategy is to identify key CG conformations and employ</span><br><span>backmapping methods to retrieve atomistic coordinates. Currently, rule-based methods often yield</span><br><span>suboptimal geometries and rely on energy relaxation, resulting in less-than-optimal outcomes. In</span><br><span>contrast, machine learning techniques offer higher accuracy but may lack transferability between</span><br><span>systems or be tied to specific CG mappings. In this study, we present HEroBM, a dynamic and scalable</span><br><span>method that utilizes deep equivariant graph neural networks and a hierarchical approach to achieve</span><br><span>high-resolution backmapping. HEroBM is capable of handling any type of CG mapping, providing a</span><br><span>versatile and efficient protocol for reconstructing atomistic structures with high accuracy. Grounded</span><br><span>in local principles, HEroBM spans the entire chemical space and can be applied across systems of</span><br><span>varying composition and sizes. We demonstrate the versatility of our framework through a range of</span><br><span>biological systems, including a complex real-case scenario. Here, our end-to-end backmapping approach</span><br><span>accurately generates atomistic coordinates for a G protein-coupled receptor bound to an organic small</span><br><span>molecule within a cholesterol/phospholipid bilayer. The high-fidelity HEroBM backmapping enables</span><br><span>researchers to effortlessly transition between CG and all-atom simulations, opening unprecedented</span><br><span>avenues for molecular investigations.</span></p>
ICELEARNING - Detection of ice core particles via deep neural networks
<p>This dataset refers to the ICELEARNING project - Detection of ice core particles via deep neural networks, by Maffezzoli N. et al., <em>The Cryosphere</em>, 10.5194/tc-17-539-2023, 2023.</p> <p>The main folder contains all TRAINING data. </p> <p>The TEST data are contained in the folder /test. </p> <p>Please refer to the <a href="https://github.com/nmaffe/icelearning">icelearning GitHub</a> repository for instructions. </p>
Guinea baboon vocalizations dataset automatically extracted with a deep neural network from natural audio recordings
<p><strong>Abstract</strong></p> <p>The data collection process consisted of continuously recording during one month a group of Guinea baboons living in semi-liberty at the CNRS primatology center in Rousset-sur-Arc (France). Two microphones we placed nearby their enclosure to continuously record the sounds produced by the group. A convolutional neural network (CNN) was used on these large and noisy audio recordings to automatically extract segments of sound containing a baboon vocal production by following the method of <a href="https://arxiv.org/abs/2302.07640">Bonafos et al. (2023)</a>. The resulting dataset consists of one-second to several-minute wav files of automatically detected vocalizations segments. The dataset thus provides a wide range of baboon vocalizations produced at all times of the day. It can be used to study vocal productions of non-human primates, their repertoire, their distribution over the day, their frequency, and their heterogeneity. In addition to the analysis of animal communication, the dataset can also be used as a learning base for sound classification models.</p> <p> </p> <p><strong>Data acquisition</strong></p> <p>The data are audio recordings of baboons. The recordings were made with a H6 Zoom recorder, using the included XYH-6 stereo microphone. The sample size is 44100 Hertz, 16 bits. The microphones were placed in the vicinity of the enclosure for one month and recorded continuously on a PC computer. A CNN passed over the data with a sliding window of 1 second and an overlap of 80% to detect the vocal productions of the baboons. The dataset consists of the segments predicted by the CNN to contain a baboon vocalization. Windows containing signal less than one second apart were merged into a single vocalization.</p> <p> </p> <p><strong>Data source location</strong></p> <ul> <li>Institution: CNRS, Primate Facility</li> <li> <p>City/Town/Region: Rousset-sur-Arc</p> </li> <li> <p>Country: France</p> </li> <li> <p>Latitude and longitude for collected samples/data: 43.47033535251509, 5.6514732876668905</p> </li> </ul> <p> </p> <p><strong>Value of the data</strong></p> <ul> <li> <p>This dataset is relatively unique in terms of the quantity of vocalizations available.</p> </li> <li> <p>This massive dataset can be very useful to two types of scientific communities: experts in primatology who study the vocal productions of non-human primates, and experts in data science and audio signal processing.</p> </li> <li> <p>The machine learning research community has at its disposal a database of several dozen hours of animal vocalizations, which will make it possible to build up a large learning base, very useful for Environemental Sound Recognition tasks, for example.</p> </li> </ul> <p> </p> <p><strong>Objective</strong></p> <p>This dataset is a follow-up of two studies on the vocal productions of Guinea baboons (Papio papio) in which we carried out analyses of their vocal productions on the basis of a relatively large vocalization sample containing around 1300 vocalizations (<a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0169321">Boë, Berthommier, Legou, Captier, Kemp, Sawallis, Becker, Rey, & Fagot, 2017</a>; <a href="https://hal.science/hal-01649539">Kemp, Rey, Legou, Boë, Berthommier, Becker, & Fagot, 2017</a>). The aim was to collect a larger database using the technique of deep convolutional neural networks in order to 1) automatically detect vocal productions in a large continuous audio recording and 2) perform a categorization of these vocalizations on a more massive sample. A description of the pipeline that enabled these automatic detections and categorizations is given in <a href="https://arxiv.org/abs/2302.07640">Bonafos, Pudlo, Freyermuth, Legou, Fagot, Tronçon, & Rey (2023)</a>.</p> <p> </p> <p><strong>Data description</strong></p> <p>The data is a set of audio files in wav format. They are at least one second long (the size of the window), up to several minutes, if several windows are consecutively predicted as containing signal. Moreover, we add the labeled data we used to train the CNN which did the prediction. We also provide two hours of the continuous recordings to have an idea of the continuous recordings and test the code of the paper provided on <a href="https://gitlab.com/papers4375727/detection-and-classification-of-vocal-productions">gitlab</a>.</p> <p>In addition, there is a database in csv format listing all the vocalizations, the day and time of their production, and the prediction probabilities of the model.</p> <p> </p> <p><strong>Experimental design, materials and methods</strong></p> <p>The original recordings represent one month of continuous audio recording. Seven hours of this month were manually labelled. They were segmented and labelled according to whether or not there was a monkey vocalization (i.e., noise or vocalization) and, if there was a vocalization, according to the type of vocalization (6 possible classes: bark, copulation grunt, grunt, scream, yak, wahoo). These manually labelled data were used as a training set for a CNN, which was automatically trained following the pipeline of Bonafos et al. (2023). This model was then used to automatically detect and classify vocalization during the whole month of audio recording. It processes the data in the same way when predicting new data as it does when training. It uses a sliding window of one second with an overlap of 80%. It does not take into account information from previous predictions, but calculates the probability of a vocalization in each one-second window independently. It then iterates through the month. For each window, the model predicts two outputs: the probability that there is a vocalization and the probability of each class of vocalization.</p> <p>For the purpose of generating the wav files, if a window has a probability of a vocalization greater than 0.5, it is considered to contain a vocalization. If it is the first one, a vocalization is started at that moment. If the time windows that follow a vocalization also contain a vocalization, then the signal they contain is added to the first segment for which a vocalization has been detected. As soon as a one-second segment no longer contains a signal corresponding to a vocalization, the wav file is closed. If windows are predicted to contain no vocalizations, but are between two windows that contain vocalizations within 1 second of each other, then all windows are merged.</p>
scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data
<p>This repository contains the training data and source code to reproduce the results of our paper:<br>scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data</p> <p>More description can be also found in GitHub (https://github.com/LPH-BIG/scGraph2Vec).</p>
Deep neural networks and humans both benefit from compositional language structure
<p>This dataset holds the results generated in the paper:</p> <p>Deep neural networks and humans both benefit from compositional language structure</p> <p>by L. Galke, Y. Ram, and L. Raviv.</p>
Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network
<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>
A Deep Neural Network Based SMAP Soil Moisture Product
<p>The soil moisture datasets here are based on a deep neural network (DNN) that utilizes the merits of a suite of existing satellite and reanalysis products to produce a new SM product with minimum (maximum) bias (correlation) -- using NASA’s Soil Moisture Active Passive (SMAP) data and ERA5 reanalysis. The benchmark of the network is a bias-adjusted SM with maximum correlation with in situ data over each land-cover type. The bias is adjusted to the product that exhibits a minimum bias over each land-cover type. Consistent with the laws of L-band microwave propagation in soil and canopy, the input variables include polarized SMAP brightness temperatures, incidence angles, vegetation scattering albedo, surface roughness parameter, surface water fraction, effective soil temperatures, bulk density, clay fraction, and vegetation optical depth from the normalized difference vegetation index (NDVI) climatology. The DNN is trained and validated using two years (04/2015--03/2017) of global data and deployed for assessment of its performance from 04/2017 to 03/2021. The testing results against in situ measurements demonstrate that the DNN outputs typically exhibit improved error quality metrics over most land cover types and climate regimes and can properly capture SM temporal dynamics, beyond each SMAP product across regional to continental scales.</p>
Prediction of inter-chain distance maps of protein complexes with 2D attention-based deep neural networks
<p>Residue-residue distance information is useful for predicting tertiary structures of protein monomers or quaternary structures of protein complexes. Many deep learning methods have been developed to predict intra-chain residue-residue distances of monomers accurately, but few methods can accurately predict inter-chain residue-residue distances of complexes. We develop a deep learning method CDPred (i.e., Complex Distance Prediction) based on the 2D attention-powered residual network to address the gap. Tested on two homodimer datasets, CDPred achieves the precision of 60.94% and 42.93% for top L/5 inter-chain contact predictions (L: length of the monomer in homodimer), respectively, substantially higher than DeepHomo’s 37.40% and 23.08% and GLINTER’s 48.09% and 36.74%. Tested on the two heterodimer datasets, the top Ls/5 inter-chain contact prediction precision (Ls: length of the shorter monomer in heterodimer) of CDPred is 47.59% and 22.87% respectively, surpassing GLINTER’s 23.24% and 13.49%. Moreover, the prediction of CDPred is complementary with that of AlphaFold2-multimer.</p>
Prediction of inter-chain distance maps of protein complexes with 2D attention-based deep neural networks
<p>Residue-residue distance information is useful for predicting tertiary structures of protein monomers or quaternary structures of protein complexes. Many deep learning methods have been developed to predict intra-chain residue-residue distances of monomers accurately, but few methods can accurately predict inter-chain residue-residue distances of complexes. We develop a deep learning method CDPred (i.e., Complex Distance Prediction) based on the 2D attention-powered residual network to address the gap. Tested on two homodimer datasets, CDPred achieves the precision of 60.94% and 42.93% for top L/5 inter-chain contact predictions (L: length of the monomer in homodimer), respectively, substantially higher than DeepHomo’s 37.40% and 23.08% and GLINTER’s 48.09% and 36.74%. Tested on the two heterodimer datasets, the top Ls/5 inter-chain contact prediction precision (Ls: length of the shorter monomer in heterodimer) of CDPred is 47.59% and 22.87% respectively, surpassing GLINTER’s 23.24% and 13.49%. Moreover, the prediction of CDPred is complementary with that of AlphaFold2-multimer.</p>
Semantic Segmentation of Time Series Imagery Using Deep Convolutional Neural Networks: A Case Study of Sandbars in Grand Canyon
<p>This dataset contains imagery used to train and test Deep Convolutional Neural Networks for the purpose of binary semantic segmentation of a time series of oblique imagery capturing sandbar monitoring sites in The Grand Canyon. In addition the scripts needed for removing image distortion, registering, rectifying, and labeling imagery is present. </p>
Dataset from the paper entitled "Complex structure of molten FLiBe (2 LiF – BeF2) examined by experimental neutron scattering, X-ray scattering, and deep neural network-based molecular dynamics"
<p>Dataset from the paper entitled "Complex structure of molten FLiBe (2 LiF – BeF2) examined by experimental neutron scattering, X-ray scattering, and deep neural network-based molecular dynamics". These data include experimental total scattering measurements and molecular dynamics simulations on the molten structure of FLiBe. </p>
AI4Life-MDC24 Challenge data: Fluorescence Microscopy Datasets for Training Deep Neural Networks
<p>This is a subset of the Supporting data for <em>Guy M Hagen, Justin Bendesky, Rosa Machado, Tram-Anh Nguyen, Tanmay Kumar, Jonathan Ventura, Fluorescence microscopy datasets for training deep neural networks, GigaScience, Volume 10, Issue 5, May 2021, giab032, <a href="https://doi.org/10.1093/gigascience/giab032">https://doi.org/10.1093/gigascience/giab032</a></em><br><br>The selected <strong>subset</strong> contains 79 images from Data Set 4 in the form of a single tiff file. <br><br>The paper describing the original dataset is available here: <a href="https://academic.oup.com/gigascience/article/10/5/giab032/6269106">https://academic.oup.com/gigascience/article/10/5/giab032/6269106</a><br>The original dataset is available here: <a href="http://gigadb.org/dataset/100888">http://gigadb.org/dataset/100888</a></p> <p><br>AI4Life has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement number 101057970. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.</p>
Automatic taxonomic identification based on the Fossil Image Dataset (>415,000 images) and deep convolutional neural networks
<p>This is a Fossil Image Dataset, which contains >415000 images. A total of 50 clades were labeled, with a final 90% accuracy. We used the web crawler to download fossil images from the Internet. We declare that all the collected images are used for academic purposes only. If anyone wants to use this dataset, please agree on the Terms of access for the Fossil Image Dataset (FID). We uploaded two datasets: FID (contains 0.415 million images) and reduced-FID (60 thousand images, 1200 for each clade). Requirements of necessary preinstalled Python libraries, algorithms for analysis, and the model weights are available at <a href="https://github.com/XiaokangLiuCUG/Fossil_Image_Dataset">https://github.com/XiaokangLiuCUG/Fossil_Image_Dataset</a>.</p>
A Deep Neural Network Based SMAP Soil Moisture Product
<p>It is demonstrated that while satellite soil moisture (SM) retrievals often have minimum biases, reanalysis data can capture more temporal variability of SM, especially for non-cropland areas -- when validated against in situ measurements. Accordingly, this paper presents a deep neural network (DNN) that utilizes the merits of a suite of existing satellite and reanalysis products to produce a new SM product with minimum (maximum) bias (correlation) -- using NASA’s Soil Moisture Active Passive (SMAP) data and ERA5 reanalysis. The benchmark of the network is a bias-adjusted SM with maximum correlation with in situ data over each land-cover type. The mean of the benchmark data is adjusted to the product that exhibits a minimum bias over each land-cover type. Consistent with the laws of L-band microwave propagation in soil and canopy, the input variables of DNN include polarized SMAP brightness temperatures, incidence angle, vegetation scattering albedo, surface roughness parameter, surface water fraction, effective soil temperatures, bulk density, clay fraction, and vegetation optical depth from the normalized difference vegetation index (NDVI) climatology. The DNN is trained and validated using two years (04/2015--03/2017) of global data and deployed for assessment of its performance from 04/2017 to 03/2021. The testing results against in situ measurements demonstrate that the DNN outputs typically exhibit improved error quality metrics over most land-cover types and climate regimes and can properly capture SM temporal dynamics, beyond each SMAP product across regional to continental scales.</p>
A tempοral Deep Convolutional Neural Network model on Sentinel-1 Image Time Series for pixel-wise Flood Classification (dataset)
<p>This is a dataset which has been designed to be used for flood time series classification. Each time series is annotated as flood or no-flood and represents a pixel-wise time series derived from stack of Sentinel-1 IW GRD images that have been pre-processed according to <a href="http://doi.org/10.5281/zenodo.6510223">https://doi.org/10.5281/zenodo.6510223</a>.</p>
DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing
<p>Curated dataset for the manuscript named "DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing".</p> <p>Five publicly available datasets sequenced on ONT MinION/GridION were used for the experiments (see below for original sources). These datasets contained raw signal data in single-FAST5 format (one file per each read), which were converted to BLOW5 format using slow5tools to enable convenient and efficient file manipulation. Then, 40,000 reads containing at least 4500 signal samples were extracted from each dataset. From each dataset, 20,000 reads are for training (<species>/train-<species>.blow5) and the rest for testing (<species>/test-<species>.blow5). Basecalled reads for the dataset used for testing are also available (test-<species>.fastq). Guppy version 6.1.3 under dna_r9.4.1_450bps_hac mode was used. The reference genomes are also given (<species>/<species>-ref.fasta)</p> <p>Original datasets are from the following sources:<br> SARS-CoV-2: https://community.artic.network/t/links-to-raw-fast5-fastq-data-for-artic-protocol/17<br> Zymo Metagenome: https://github.com/LomanLab/mockcommunity<br> Chlamydomonas: https://sra-download.ncbi.nlm.nih.gov/traces/era20/ERZ/003237/ERR3237140/Chlamydomonas_0.tar.gz<br> Saccharomyces cerevisiae: https://www.ncbi.nlm.nih.gov/bioproject/PRJNA510813</p>
Results from Interpreting Cis-Regulatory Interactions from Large-Scale Deep Neural Networks for Genomics
Open the record for dataset details and reuse information.
Figure 1. CNN architecture (adopted from Krizhevsky et al. '12)-Measuring Customer Behavior with Deep Convolutional Neural Networks
<p>The architecture of a CNN can be described as following. A small pixel region goes to input neurons and then connects to a first convolution hidden layer (Figure1). There we can see a set of learnable filters, which are activated during the presentation some particular type of feature in pixel region in the input. On this phase, CNN does shift invariance, which is carried by feature map. Subsampling layer goes next. There we have two processes: local averaging and sampling. As a result, we get declining resolution of feature map. To correspond this task CNN needs supervised learning. Before starting the experiment, we gave a set of labeled videos with different emotional experience. The system analyses images and finds similar features. Then the system creates a map, where it arranges videos in accordance with similar features. Thereby, images with similar emotions form certain class. To test the system, we add other videos and correct the system when it refers them improperly. The proposed model consists of four convolutional layers, followed by max-pooling layers, and three fully-connected layers with a final classificatory presented with MLP (with six basic outputs, corresponding to basic emotions for emotion classification and two outputs for motion classification for typical and non-typical behavior). The input data was presented as infrared camera output.</p>
gazeNet: End-to-end eye-movement event detection with deep neural networks
<p>This repository contains synthetic eye-movement dataset used to train deep learning based eye-movement event detection algorithm described in Zemblys, R., Niehorster, D.C. & Holmqvist, K. (2018). gazeNet: End-to-end eye-movement event detection with deep neural networks. Behavior research methods, pp 1–25. <a href="https://doi.org/10.3758/s13428-018-1133-5">https://doi.org/10.3758/s13428-018-1133-5</a></p> <p>Code used to train a model can be found on github <a href="https://github.com/r-zemblys/gazeNet">here</a>. Code to generate synthetic eye-movement data can be downloaded from <a href="https://github.com/r-zemblys/gazeGenNet">here</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.