Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

921

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

921 results for “neural networks”

Learn how ShareScore rates datasets ↗
zenodo32/100

Convolution, aggregation and attention based deep neural networks for accelerating simulations in mechanics [Dataset]

<p>Supplementary data for &#39;Convolution, aggregation and attention based deep neural networks for accelerating simulations in mechanics&#39;.&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Synthetic plant modelling: creating plants in 3D to train neural networks

<p><strong>The following video describes how virtual plant modeling can be used to create synthetic datasets that inform machine learning and computer vision. funded by EU Grant 773875</strong></p> <p><em>Videos are available in:</em></p> <ul> <li>Hi-res (1080p Apple ProRes)</li> <li>Mid-res&nbsp;(1080p&nbsp;H265)</li> </ul> <p><strong>Video script:</strong></p> <p>(MINCHIN) At INRIA in Lyon in France we were also looking at virtual modelling of plants. This creates synthetic data sets that can then help us train machine learning algorithms to identify plants in the field or particular traits in real life.<br> <br> (GODIN) So this is the current bottleneck where we are, we need absolutely massive ground truth data if we want to train this algorithm, this new algorithm, these new families of algorithms to produce this 3d structures and segment it in the the correct botanical way. So how to do this? There are two options basically; the first option would be to acquire a real plant data by photograph, by scanner laser, and then the expert segment by hand different parts of the plant, saying that this is a leaf, this part of the point cloud is a leaf, this part of the point cloud is a stem et cetera. And you can imagine that this is extremely time consuming, this is extremely heavy task and we are blocked at this point because of the of the ability of humans to do such complicated tasks.<br> <br> And there is another option that is to create artificial plants and then say whether with this segmentation of the virtual plant. Whether we are right or not because as we designed the virtual plant we know that this part of the point cloud corresponds to a leaf, this part of the point cloud corresponds to a stem and because of this it is possible to automatise the training of the system.<br> <br> So what we want to do in the context of such a phenotyping approach is to use a virtual pipeline where we would produce the virtual plant in the computer. Then we would create point clouds out of these virtual plants, so this would be virtual point clouds. Then we would use training algorithms in the context of this machine learning construction process, and then we would get as an output the trained machine learning system. Then once we have this, it is possible to get back to the original pipeline and use here this trained machine learning system in order to recognise identify segment the different organs on the plant.<br> <br> So L-py is the programming language to simulate plants and from this it is possible to create full databases of plants by a stochastic simulation of this, of plant populations, to produce massive data, and then this massive data, give them to machine learning systems in order to train this machine learning based on this large amount of input data.<br> <br> The first thing to know is that how plants are growing. You have leaves like this and then you have a stem like this. This is growing due to this small part of the tip that is called the apex, and the apex is producing all the organs that are being built on the plant. So for the leaves and also the fruits here or the flowers. So if you look at this in a more detailed manner you you make a close-up on this. What you observe is that it is like a small dome like this, that is producing lateral organs and this is the stem, and these are the young organs. Like this can be a leaf or a flower or anything else that is produced by the by the meristem. And this part here, it is the place where all the stem cells are living and they are dividing and they are producing the small organs here one after the other at the tip of the plant. So if we want to model this we need to model how this small part of the plant which is built, which is made up of a small amount of stem cells, undifferentiated cells, how this is growing. an apex let&#39;s call it &lsquo;a producing a piece of stem&rsquo;, (that I call for the internet) and literally produces an apex that will in turn be able to grow on it on its own. We formalise this rule of growth by this, let&#39;s say, a mathematical expression.<br> <br> (BESNARD) In the case of ROMI, so this European project we are involved in, we are using virtual plants for a precise objective which is to use virtual plants to create a data set, virtual dataset that we can use for a training deep learning algorithm very efficiently and costless, and as I told you. Then, once you have an objective you have to question yourself whether realism and the realistic rendering of the plant is useful or not for your objective. In the program. We use pipelines that are able to detect automatically to segment the plant and to detect automatically the organs, okay so here basically it will be the branching point that will be discovered and this branching point can be either branches that are cut here or the helix here. And once it is detected we use a representation where we highlight these branching points so this or branches, and you can see that here you have another layer and this highlight overlays perfectly with the branches. So you can say wow it&#39;s really good, but when you dig into these plants sometimes algorithms, machine learning algorithms are not working well and here is when you proceed with the same algorithm, the leaf is recognised partly as a leaf at the beginning okay, but the tip of the leaf is recognised as a fruit or a branch. So it shows you here, that you have some issues so those machine learning algorithms are not perfect and we need to improve them to increase our performances in terms of fragmentation.<br> <br> So I will proceed in the natural plant, and you can see that definitely the the leaves here are in the current model with which we trained this network, they are like quiet simple structure so it&#39;s a flat very simple shape that are not like really the shape of the current leaf. So here&#39;s an important question if we make a new plant modified here leaf shapes, would this help in this the algorithm to make a better plan segmentation?<br> <br> (GODIN) To make more realistic plants we decided to grow real Arabidopsis plants in growth chamber and to measure them in a systematic manner. We then used these detailed measurements to refine our virtual Arabidopsis plant in different ways. First we refined the modelling of the different organs for example the coiling of leaves. With time the leaf would be produced laterally by the stem in a sort of straight way like this, or like this, and then with time it would curl and change the shape. Instead of having one single curve. Now I will have a set of curves, I would just put the curve, it would be straight in the beginning then i would fold a bit the curve fall a bit the curve fold a bit like this, even I can go in this way. So I would define the curve, zero curve, one curve, six. Okay actually this is what I did here. And then I say all this bunch, create me an object that is able, given a time . So the Curve, Function, Object, &lsquo;curve function object&rsquo;, it is able to take a time too &lsquo;Tao&rsquo; (T) and then to return a curve object as at the time term. And then so this curve object is provided by L-pi. You just call it with a series of functions, then once you have this you can call curve of &rsquo;T&rsquo; and then if I know this, I can compute the current curve at time, and then display it and and actually while it is doing this, it is able to compute all the intermediate curves by interpolation. Like this.<br> <br> Detailed models of organs were then assembled to model groups of organs, such as flowers or at an even more integrated level the dynamics of inferences development. The difficulty here is to synchronise the growth dynamics of each part, for this we use the new strategy that we developed in the course of the project and that is called &lsquo;hierarchical timeline warping&rsquo;. This strategy makes it possible to synchronise the growth dynamics of the different plant parts onto each other in a hierarchical and non-linear manner similarly still based on real plant measurements. We modelled also the observed variability in the dimensions and orientation of the different organs. The gravitropism and the mechano-perception of the different axes resulting in complex and dynamic bending of their parts. The final model provides a realistic rendering of the plant growth that can be used as a faithful reference to train ROMI&rsquo;s machine learning algorithms.<br> <br> Then the technology to construct virtual plants developed for Arabidopsis was used to generate other virtual plants with different levels of complexity and accuracy, we first updated our tomato model by introducing stochasticity in the development of the plant here you can observe two tomato plants that were generated using the same stochastic model. The model varies the number of organs their size their dimensions their orientation with respect to their parent stem, their bending and so on. Another virtual model was made of a Canopodium this plant is a weed that can be easily found in crop fields similarly to Arabidopsis although with less details, we grew several plants in growth chamber and observed how it grows. We used these observations to construct a virtual model model of a Canopodium, the model is very different from that of a tomato, it grows with a main axis that dominates over secondary axis that grows in turn with slight delay with respect to the leader. Here also stochasticity was introduced in the model so that one can generate different individuals with a simple click.<br> <br> Finally based on the training course participants could start to produce their own virtual plants. Here is the first version of a model of a pepper plant showing the step-by-step approach of the student based on real plant observations. Here is a 3d model of a carrot plant with fractal leaves and stochasticity. Several instances of this stochastic model can be generated to create a virtual field of carrots. So by the way we will learn on Friday, Ayan will make a presentation showing how we can use these systems and to scan them make a sample of points, a 3d cloud sample of points, so that we can train systems, machine learning systems, that take as an input sample points, in order to produce the right segmentation of the point cloud.<br> <br> Parts of this video were extracted from a week-long course given in April 2022 and covering every aspect of how to virtually model plants in three dimensions, with L systems in Python. If you are interested don&#39;t hesitate to go and check it out on the ROMI youtube channel.</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

Video simulations for paper "Rapid Spatio-Temporal Flood Modelling via Hydraulics-Based Graph Neural Networks"

<p>Videos of the comparison between numerical and deep learning simulations for test datasets 1, 2, and 3 for paper &quot;Rapid Spatio-Temporal Flood Modelling via Hydraulics-Based Graph Neural Networks&quot;.</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

Data: Upsamling Monte Carlo Neutron Transport Simulation Tallies Using a Convolutional Neural Network

<p>This repository contains:</p> <ul> <li>openmc-data-XXXX.tar.gz - Training Data generated with the OpenMC Monte Carlo code representing neutron flux tallies in 4,400 unique light water reactor fuel assemblies in HDF5 format. Training samples consist of tallies in 64x64 pixels and 8 neutron energy groups, and tallies in 128x128 pixels and 16 neutron energy groups. Folders 0008 to 0023 contain training and validation data. Folder 0024 contains test data.</li> <li>out.mat - Upsampling results using a Convolutional Neural Network for 300 testing data samples in MATLAB format. These data include OpenMC tally uncertainties in low and high resolution tallies, scaling values used in data pre-processing, low resolution inputs to the CNN, and high resolution upsampled results as well as high resolution ground truth values.</li> </ul>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Artificial neural network model and metabolomics data of selected microbial strains

<p>Metabolomics data, metadata, sample R code, and a pre-trained artificial neural network model to predict group memberships of the bacterial strains in the dataset.</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Multivariate prediction on wake-affected wind turbines using graph neural networks (Eurodyn) database

<p>Database consisting of graphs generated using randomized layouts and PyWake simulations used in&nbsp;&#39;<em>Multivariate prediction on wake-affected wind turbines using graph neural networks</em>&#39;, contribution to Eurodyn 2023.&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Raw datasets for paper "Rapid Spatio-Temporal Flood Modelling via Hydraulics-Based Graph Neural Networks"

<p>Raw datasets for paper &quot;Rapid Spatio-Temporal Flood Modelling via Hydraulics-Based Graph Neural Networks&quot;.</p> <p>The zip folder comprises 4 subfolders (DEM, WD, VX, VY), containing the elevation, water depths in time, and velocities (in x and y directions) in time for all training and testing simulations. The overview.csv file provides the runtime of the numerical model on each different simulation, identified by its id.</p> <p>The simulations ids are divided as follows:</p> <p>- 1-80: Training and validation</p> <p>- 501-520: Testing dataset 1</p> <p>- 10001-10020: Testing dataset 2</p> <p>- 15001-15020: Testing dataset 3</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

Multi-scale full waveform inversion based on a convolutional neural network

<p>The research data from this paper are uploaded here and are available for download.</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Improving rainfall forecast at the district scale over the eastern Indian region using deep neural network

<p>The code contains ANN and CNN models.</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Single-Shot Optical Neural Network

<p>Data from classification of MNIST [1], Fashion-MNIST [2] and QuickDraw [3] images reported in&nbsp;&quot;Single-Shot Optical Neural Network&quot; by L. Bernstein et al. The networks of 784 -&gt; N (-&gt; N) -&gt; 10 activations performed inference on the test sets through consecutive matrix products implemented on the optical hardware, with ReLU applied electronically between each layer (see main text for more details). Folders for each tested network contain the following text files:</p> <ul> <li>Inputs: 2D matrices of size B&nbsp;x 784&nbsp;containing training (B = 50,000 or 100,000), validation (B = 10,000) and test (B = 10,000) sets. Each&nbsp;row is an input vector that can be reshaped&nbsp;into an input image&nbsp;of size 28 x 28.</li> <li>True labels: B-length vectors containing the true label of each input in&nbsp;the training, validation and test sets.</li> <li>Neural network weights: 2D matrices of size K x N&nbsp;used in inference experiments to classify the test sets. Weights were pre-trained on the training set&nbsp;using a digital electronic computer as described in Materials and Methods. Weight values were normalized such that all values fall between -255 and 255. In the optical neural network, the weighting SLM displays&nbsp;the absolute values of the weights (rounded to the nearest integer), and the negative weight signs are applied in post-processing. The&nbsp;32-bit weight values were used for inference performed on the&nbsp;digital electronic computer for the ground truth comparison.</li> <li>Outputs (normalized):&nbsp;2D output matrices of size 10,000 x 10&nbsp;from the networks processed on a digital electronic computer (ground truth) and the&nbsp;optical neural network. Each row is an output vector where the position of the&nbsp;maximum value indicates the predicted label of the input in the same row of the test set.</li> <li>Predicted labels: 10,000-element vectors that represent the labels predicted by the networks processed on a digital electronic computer (ground truth) and the&nbsp;optical neural network. These predicted labels were used to generate the confusion matrices and calculate the classification accuracies (versus the true labels).</li> </ul> <p>The classes for the Fashion-MNIST dataset are the following:</p> <ul> <li>0: T-shirt</li> <li>1: Trouser</li> <li>2: Pullover</li> <li>3: Dress</li> <li>4: Coat</li> <li>5: Sandal</li> <li>6: Shirt</li> <li>7: Sneaker</li> <li>8: Bag</li> <li>9: Ankle boot</li> </ul> <p>And the randomly selected classes for QuickDraw are:</p> <ul> <li>0: Hourglass</li> <li>1: Saw</li> <li>2: Golf club</li> <li>3: See saw</li> <li>4: Spoon</li> <li>5: Horse</li> <li>6: Onion</li> <li>7: Light bulb</li> <li>8: Harp</li> <li>9: Flip flops</li> </ul> <p>[1]&nbsp;Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, Gradient-based learning applied to document recognition. <em>Proceedings of the IEEE</em> <strong>86</strong>, 2278&ndash;2324 (1998).</p> <p>[2]&nbsp;H. Xiao, K. Rasul, R. Vollgraf, Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms. Preprint at https://arxiv.org/abs/1708.07747 (2017).</p> <p>[3] J. Jongejan, H. Rowley, T. Kawashima, J. Kim, N. Fox-Gieg, The Quick, Draw! AI experiment, https://quickdraw.withgoogle.com/ (2016).</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Uncovering Stress Fields and Defects Distributions in Graphene Using Deep Neural Networks

<p>The trained neural networks, complete data set, and MATLAB script used to generate molecular dynamics simulation files are available here.</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Learning a quantum computer's capability using convolutional neural networks

<p>This is supplemental data and code for: D. Hothem et al., <em>Learning a quantum computer&#39;s capability using convolutional neural networks, </em>(to be published).</p> <p>This folder contains all the data and the analysis code to generate the results presented in that paper. The core data analysis routines use PyGSTi, which can be found at&nbsp;<a href="https://github.com/pyGSTio/pyGSTi">https://github.com/pyGSTio/pyGSTi</a>.</p> <p>Please direct any questions to Daniel Hothem (dhothem@sandia.gov).</p> <p>NOTE: This description template was borrowed from Timothy Proctor&#39;s Zenodo entry for: Scalable Randomized Benchmarking of Quantum Computers using Mirror Circuits.</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Supplementary dataset for paper: "Approximate non-linear model predictive control with safety-augmented neural networks"

<p>Supplementary dataset for paper Henrik Hose and Johannes Koehler and Melanie N. Zeilinger and Sebastian Trimpe &quot;Approximate non-linear model predictive control with safety-augmented neural networks&quot;.</p> <p>The code to use this dataset is publicly available at&nbsp;<a href="https://github.com/hshose/soeampc">https://github.com/hshose/soeampc</a></p> <p>The dataset contains training and testing data to train an NN controller for three standard benchmark systems, a stir tank reactor, a quadcopter, and a chain mass system.</p> <p>For each system, there are initial conditions as comma separated value in the `x0.txt` file, the MPC input trajectory in the `U.txt` file and the corresponding predicted state sequence in the `X.txt` file. MPC parameters are provided for each system. The dataset was computed using acados for SQP solving.</p> <p>The dataset also contains pretrained neural network approximations of the dataset.These are provided in the `pretrained_models.zip` file. The neural networks were trained with tensorflow.</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Ensemble Ecological Niche Models, in 2019 and across RCP 2.6, 4.5, and 8.5 scenarios in 2050 and 2100, of 1508 European Marine Species based on Ecological Niche Models developed with Artificial Neural Networks, Maximum Entropy, Support Vector Machines, and AquaMaps at 0.5° Resolution

<p>Ensemble Ecological Niche Models, in 2019 and across RCP 2.6, 4.5, and 8.5 scenarios in 2050 and 2100, of 1508 European marine species based on Ecological Niche Models developed with (i) Artificial Neural Networks, (ii) Maximum Entropy, (iii) Support Vector Machines, and (iv) AquaMaps at 0.5&deg; Resolution. The data report, for each 0.5&deg; cell, how many models (from 0 to 4) overcome a model-specific decision threshold to assess species presence in the cell.</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Maximally Error-Tolerant MZI Mesh-based Optical Neural Networks

<p>Codes.zip contains code that trains and tests maximally error-tolerant Mach-Zehnder mesh optical neural networks on MNIST, FashionMNIST, and KMNIST.</p> <p>models.zip contains all the trained model parameters.</p>

opencc-by-4.0May 2023View details →
zenodo32/100

Application of deep neural networks to reconstruct coastal water quality data

<p>In this study ordinary and new integrated deep neural networks were developed for the reconstruction of measured specific conductance (<em>SC</em>) data. Five stations of USGS in the Gulf of Mexico were considered as case study.</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Interpreting Cis-Regulatory Interactions from Large-Scale Deep Neural Networks for Genomics

<p>Results and code to replicate analysis in &quot;Interpreting Cis-Regulatory Interactions from<br> Large-Scale Deep Neural Networks for Genomics&quot; by Toneyan and Koo.</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Dataset for: Star Photometry for DECaLS and SDSS Images Based on Convolutional Neural Networks

<p>The dataset consists of two parts, the training dataset and the comparison dataset. The training dataset and the comparison dataset are each divided into three parts, namely the simulation dataset, the SDSS dataset and the DECaLs dataset.</p> <p>The simulation dataset is simulated by PhoSim software and the original size of the simulation data is 512x512 pixels. The SDSS dataset is from DR12 and the ObjId of the target as well as other parameters are given in the csv file.The DECaLs dataset is from DR9.</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

PlayNet: Real-time Handball Play Classification with Kalman Embeddings and Neural Networks Dataset

<p>Handball play classification dataset. On one hand we have player position and direction (x, y, vx, vy)&nbsp;estimated with a Kalman Filter, and on the other the ball position. For each tuple there is an associated play&nbsp;(right_attack, left_attack, right_transition, left_transition, time_out, right_penal, left_penal). There are two splits train and test.</p> <p>&nbsp;</p> <p>&nbsp;</p>

openother-ncOct 2022View details →
zenodo32/100

Dataset and code for the manuscript 'Parameterizing Vertical Mixing Coefficients in the Ocean Surface Boundary Layer using Neural Networks'

<p>This repository contains the code and data used in the manuscript&nbsp;&#39;Parameterizing Vertical Mixing Coefficients in the Ocean Surface Boundary Layer using Neural Networks&#39;.&nbsp;<br> Manuscript authors: Dr. Aakash Sane, Dr. Brandon G. Reichl, Dr. Alistair Adcroft, Dr. Laure Zanna<br> Manuscript preprint link:&nbsp;https://doi.org/10.48550/arXiv.2306.09045</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record