Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
921
datasets available to search
ShareScore release 0.7.1
Dataset results
921 results for “neural networks”
Surfaces/regoliths used in the training and testing of the deep neural network for surface reconstruction from simulated exospheric measurements
<p>This dataset contains the surfaces/regoliths in terms of elemental surface composition used in v2.0 - v2.5 of the paper collection: "Conceptual framework for the application of deep neural networks to surface composition reconstruction from Mercury’s exosphere".</p>
Inputs and outputs for exospheric simulations used in the deep neural network for surface reconstruction from simulated exospheric measurements
<p>This dataset contains the inputs and outputs of the exospheric simulations performed for v2.0 - v2.5 of the paper collection: "Conceptual framework for the application of deep neural networks to surface composition reconstruction from Mercury’s exosphere".</p>
Inputs and outputs for the training and testing of a deep neural network for surface reconstruction from simulated exospheric measurements
<p>This dataset contains inputs (datasets) and outputs (trainings and tests) used in v2.0 - v2.5 of the paper collection: "Conceptual framework for the application of deep neural networks to surface composition reconstruction from Mercury’s exosphere".</p>
Deep learning based on convolutional neural networks to classify nanobiomechanical data
Open the record for dataset details and reuse information.
Fast and Accurate Prediction of Tautomer Ratios in Aqueous Solution via a Siamese Neural Network
<p>Tautomerization plays a critical role in many chemical and biological processes, impacting the molecular stability, reactivity, biological activity, and ADME-Tox properties. Many drug-like molecules exist in multiple tautomeric states in aqueous solutions and complicating drug discovery. Predicting these tautomeric ratios and identifying the predominant species rapidly and accurately is crucial for computational drug discovery. In this study, we introduce sPhysNet-Taut, a deep learning model fine-tuned with experimental data leveraging the Siamese network. This model predicts tautomer ratios in aqueous solution using MMFF94-optimized geometries directly. On an experimental test set, sPhysNet-Taut surpasses all other methods, achieving state-of-the-art performance with an RMSE of 1.9 kcal/mol on the 100-tautomers set and an RMSE of 1.0 kcal/mol on the SAMPL2 challenge, and providing the best ranking power for tautomer pairs. Additionally, our results demonstrate that fine-tuning on experimental data significantly improves model performance compared to training from scratch. This work not only provides a useful deep learning model for predicting tautomer ratios, but also provides a protocol for modeling pairwise data. To facilitate user-friendliness, we developed a readily accessible tool to predict stable tautomeric states in aqueous solutions, enumerating all possible tautomeric states and ranking them using the sPhysNet-Taut model.</p>
Pretraining convolutional neural networks for mudstones petrographic thin section image classification
<p>This dataset was used in the paper "Pretraining convolutional neural networks for mudstones petrographic thin section image classification"</p>
Artificial Neural Network Symbol Demapper for Coherent Optical Fiber Systems
<p>M-files and datasets that implement an artificial neural network (ANN) demapper targeted to the compensation of fiber nonlinearities in coherent optical transmission systems. </p> <p>The dataset contains simulation data of a 11-channel WDM fiber link with numerical propagation implemented by the split-step Fourier method over standard single-mode fiber with 100 km per span and inline optical amplification with 5 dB noise figure. The launched optical power is varied in the range of 0 to 5 dBm and the distance is swept up to 30 fiber spans. The transmitted signal is a root-raised cosine single-carrier 16QAM at 64 Gbaud. </p>
Supplementary Information for Consonance-emerging Hebbian Learning neural network model predicts discreteness of musical scales and the Natural Just Intonation scale
<p><strong>The following phenomena and features are apparent in music and auditory perception in general: the discreteness of the tones in musical scales</strong> [1]<strong>, the prevalence of the tonal frequency span of one semitone (100 cents) in musical scales across cultures </strong>[1]<strong>, the list of tonal intervals ordered by consonance [2], and the musical performers’ preference of the Natural Just-Intonation scale [3] (A). However, researchers still have no agreement about the causes and the emergence of said phenomena (A). Here we show that the consonance-pattern emerging neural network model introduced in our previous study [4], predicts and yields all the said phenomena (A) with a precision of 1/100<sup>th</sup> of a semitone (1 cent). This precision is beyond the resolution of human hearing </strong>[5], [6], [7]. <strong>Since the Hebbian learning paradigm and harmonicity are the main features of our model, we propose that they are sufficient conditions for any system to yield the said phenomena (A). Therefore, they have a crucial role in processing pitch, consonance, and music perception in general. As a consequence, we additionally propose that the mentioned phenomena (A) are a balanced result of the joint workings of the Hebbian paradigm (nurture and cultural exposure) and harmonicity (auditory physics and biology).</strong></p>
Relativistic electron model in the outer radiation belt using a neural network approach
<p>This dataset includes the models, dataset, and extra figures for the paper titled</p> <p>Relativistic electron model in the outer radiation belt using a neural network approach</p> <p> </p>
Dataset - seismic data from central-western Italy used in the paper on rapid prediction of ground motion using a Convolutional Neural Network
<p>The dataset published here is the central-western Italy dataset used in the paper "<em>Transfer learning: Improving neural network based prediction of earthquake ground shaking for an area with insufficient training data"</em> (<a href="https://arxiv.org/abs/2105.05075">https://arxiv.org/abs/2105.05075</a>). The code for the paper is available at <a href="https://github.com/djozinovi/TLpredIM">https://github.com/djozinovi/TLpredIM</a>. The abstract of the paper:</p> <blockquote> <p>In a recent study (Jozinović et al, 2020) we showed that convolutional neural networks (CNNs) applied to network seismic traces can be used for rapid prediction of earthquake peak ground motion intensity measures (IMs) at distant stations using only recordings from stations near the epicenter. The predictions are made without any previous knowledge concerning the earthquake location and magnitude. This approach differs from the standard procedure adopted by earthquake early warning systems (EEWSs) that rely on location and magnitude information. In the previous study, we used 10 s, raw, multistation waveforms for the 2016 earthquake sequence in central Italy for 915 events (CI dataset). The CI dataset has a large number of spatially concentrated earthquakes and a dense station network. In this work, we applied the CNN model to an area around area near Pisa, Italy. In our initial application of the technique, we used a dataset consisting of 266 earthquakes recorded by 39 stations. We found that the CNN model trained using this smaller dataset performed worse compared to the results presented in the original study by Jozinović et al. (2020). To counter the lack of data, we adopted transfer learning (TL) using two approaches: first, by using a pre-trained model built on the CI dataset and, next, by using a pre-trained model built on a different (seismological) problem that has a larger dataset available for training. We show that the use of TL improves the results in terms of outliers, bias, and variability of the residuals between predicted and true IMs values. We also demonstrate that adding knowledge of station positions as an additional layer in the neural network improves the results. The possible use for EEW is demonstrated by the times for the warnings that would be received at the station PII.</p> </blockquote>
[Re] An anatomically constrained neural network model of fear conditioning
<p>The results contained within this archive correspond to the Python re-implementation of the computation model and replication of the classical conditioning experiment described in Armony et al. (1995). The data was generated by running the program using the 14th frequency as the Conditioned Stimulus (CS_IDX = 13) and setting the random seed for the <a href="https://numpy.org/">Numpy</a> library to 3 (NUMPY_SEED = 3).<br> During the pre- and post-conditioning testing phases, the activation values of all the neurons in the model have been recorded in different <a href="https://pandas.pydata.org/">pandas.DataFrames</a>. At the end of the experiment, those DataFrames have been written to disk using the <a href="https://hdfgroup.org/">HDF5</a> file format. It should be noted that although this might have been unnecessary given the size of the final dataset, the file has been further compressed to save on space.</p> <p>The HDF5 file format works similarly to dictionaries in Python, or Maps in other programming languages. That is, the data is organized into tables/arrays each associated with a unique key. In the case of the current dataset, the keys are the name of the different layers in lowercase (i.e.: mgm, mgv, cortex, and amygdala). Then, the array corresponding to each of those key includes the layer's neural activities for all frequencies, and for both the pre- and post-conditioning phases.<br> The columns making up each table are:</p> <ul> <li>The "Frequency" index with values in the range [1-15],</li> <li>One column for storing the activity of each unit ("Unit 1", ..., "Unit N", where N = 3 or N = 8 depending on the layer),</li> <li>The last column, entitled "Phase", contains string representations of the phase during which the activity was recorded (either "Pre-conditioning" or "Post-conditioning").</li> </ul> <p>The data included in the archive can be retrieved and stored in a dictionary for further processing using the following Python script:</p> <pre><code class="language-python">import pandas as pd # DATA_PATH is the absolute path to the file containing the hdf5 formated data data = {k: pd.read_hdf(DATA_PATH, key=k) for k in ['mgm', 'mgv', 'cortex', 'amygdala']}</code></pre> <p> </p>
Dataset for: Synthetic Micrographs of Bacteria (SyMBac) Allows Accurate Segmentation of Bacterial Cells Using Deep Neural Networks
<p>Datasets for the paper Synthetic Micrographs of Bacteria (SyMBac) Allows Accurate Segmentation of Bacterial Cells Using Deep Neural Networks, published in BMC Biology.</p>
Dataset for "Invariance of Object Detection in Untrained Deep Neural Networks"
<p><strong>Dataset for<br> "Invariance of Object Detection in Untrained Deep Neural Networks"</strong><br> Jeonghwan Cheon, Seungdae Baek, and Se-Bum Paik*<br> *Contact: sbpaik@kaist.ac.kr<br> <br> To run demo codes for "<a href="https://github.com/vsnnlab/Invariance">Invariance of Object Detection in Untrained Deep Neural Networks</a>", please download files below.<br> <br> <strong>1. Image.zip</strong><br> <strong>- Object dataset (Foldername: selectivity_var)</strong>: This set was used to find units that selectively respond to a specific object class. It contains nine object classes (bed, chair, desk, dresser, nightstand, monitor, sofa, table, toilet) and 200 images are prepared to an object class. Each image has different object identities, which means it rendered from different object 3D models (Princeton ModelNet, a 3D CAD model dataset for computer vision and cognitive science [https://modelnet.cs.princeton.edu/]). To render image of object dataset, horizontal viewpoint variation angle was randomly set between -30° and +30°. In object dataset, brightness and contrast of images are statistically comparable across the object class.<br> <strong>- Viewpoint dataset for invariance test (Folder name: invariance_test)</strong>: This set was used to test the viewpoint invariant characteristic of object selective units. This dataset consists of 13 subsets which has different viewpoints from -180° to +180° in linear scale step. It contains 200 different object identities in an object class, which are the same as those used in the object dataset.<br> <strong>- Viewpoint dataset for finding invariant unit (Folder name: invariance_unit)</strong>: This set was used to find object selective units that specifically or invariantly responded to object images of different viewpoints. This dataset consists of five angle-based viewpoint classes (-60°, -30°, 0°, 30°, 60°) with 50 object identities which were not used to find object selective unit<br> <strong>- SVM dataset (Folder name: SVM_var)</strong>: This set was used to train and test SVM which performs object detection task. It contains 60 different object identities in an object class, which were not used to find object selective unit. Specifically, it consists of 18 subsets which has different viewpoint variation range from 0° to 180°. For example, subset with 180° viewpoint variation range contains images which shows different viewpoints of objects within range of -90° and +90°.</p>
Displacement measurement via self mixing interferometry and neural network training set
<p>Self mixing interferometry is a simple and robust sensing method which can be used (among other things) to measure the displacement of a target along the light propagation axis. While conceptually simple, the actual use of this method is less straightforward than originally envisioned because reconstructing the target displacement from the interferometric signal is often tricky. A small neural network can do this task very well after proper training, as described in [10.1364/OE.419844]. This data set was used to train the network in that work (after data augmentation). It consists of a python dictionary with two keys: `truth` and `signal`. The `truth` part is a 195011-elements long numpy array corresponding to the displacement of the target in units of wavelength per 1.024 ms. The `signal` part is the interferometric signal corresponding to the displacement. It is arranged in a (195011,256,1) numpy array. Each segment of length 256 corresponds to the interferometric signal acquired during a 1.024 ms time window. For instance, the displacement value in `truth[618]` corresponds to the interferometric signal segment `signal[618,:,0]`.</p>
An Improved Tandem Neural Network Architecture for Inverse Modeling of Multicomponent Reactive Transport in Porous Media
<p>This data includes the training and testing dataset for DNN design and the observation data of synthetic example for validation. </p> <p>The code of TNNA-AUS inversion method.</p>
Test dataset for "Rapid estimation of cortical neuron activation thresholds by transcranial magnetic stimulation using convolutional neural networks"
<p>Data corresponding to test dataset used in Aberra AS, Lopez A, Grill WM, Peterchev AV. (2022). "Rapid estimation of cortical neuron activation thresholds by transcranial magnetic stimulation using convolutional neural networks". bioRxiv. Dataset includes:</p> <ul> <li><em>simnibs/ -</em> SimNIBS mesh and E-field solution file used in test dataset (posterior-anterior TMS of M1 in <em>ernie</em> example mesh, meshed with mri2mesh pipeline)</li> <li><em>layer_data/ - </em>surface meshes used for placing and orienting neuron models and corresponding sampling grids for CNNs</li> <li><em>nrn_sim_data/ - </em>Thresholds from NEURON simulations for all 25 model neurons included in the study, each at 4,999-5,000 positions and 12 azimuthal orientations ("ground truth" for CNN) </li> <li><em>cell_data/</em> - Coordinates and morphology information for all model neurons</li> <li><em>weights/</em> - Trained 3D convolutional neural networks for estimating neuron model-specific TMS thresholds given input E-field distributions on a 3D grid (see code/manuscript for dimensions)</li> <li><em>est_data/ </em>- Output of trained CNNs on all E-field data for test dataset <em> </em></li> </ul> <p> </p>
Code and extensive data for training neural networks for radiation, used in "Implementation of a machine-learned gas optics parameterization in the ECMWF Integrated Forecasting System: RRTMGP-NN 2.0""
<p>Data and code used in a paper submitted to JAMES titled :<em> Implementation of a machine-learned gas optics parameterization in the ECMWF Integrated Forecasting System</em></p> <p>1) The files <strong>ml_training_*.7z</strong> contain extensive datasets (in NetCDF format) for training neural network versions of the RRTMGP gas optics scheme as described in the paper. The datasets are read by <a href="https://github.com/peterukk/rte-rrtmgp-nn/blob/main/examples/rrtmgp-nn-training/ml_train.py">ml_train.py.</a></p> <p>2) The ML datasets were in turn generated using the input profiles (in NetCDF format) inside <strong>inputs_to_RRTMGP.zip </strong>by running the Fortran programs <code>rrtmgp_sw_gendata_rfmipstyle.F90 and rrtmgp_lw_gendata_rfmipstyle.F90 </code>in <em>rte-rrtmgp-nn/examples/rrtmgp-nn-training</em>, which call the RRTMGP gas optics scheme, The input profiles contain <strong>millions of columns, hundreds of perturbation experiments (including hypercube-sampled gas concentrations), are derived from several different data sources (including CAMS reanalysis, GCM, and CKDMIP-MMM), and span present-day, preindustrial, and future atmospheric conditions.</strong> They could be used to generate training data for developing emulators of the full RTE+RRTMGP radiation scheme, not just gas optics (see nn_dev on the <a href="https://github.com/peterukk/rte-rrtmgp-nn">RTE+RRTMGP-NN repository on Github</a>, used in a previous paper where different emulation methods were compared)</p> <p>3) The Fortran and Python code used for data generation and NN training are found in<a href="https://github.com/peterukk/rte-rrtmgp-nn/tree/main/examples/rrtmgp-nn-training"> <em>rte-rrtmgp-nn/examples/rrtmgp-nn-training</em> </a>on the main branch on Github; <strong>an archived version is also included here </strong>(<strong>rte-rrtmgp-nn-2.0.zip</strong>). See the readme in the above sub-directory for further information.</p> <p> </p>
Data for: Reconstructing Cosmological Initial Conditions from Late-Time Structure with Convolutional Neural Networks
<p>Trained models and evaluation data for the revised submitted paper "Reconstructing Cosmological Initial Conditions from Late-Time Structure with Convolutional Neural Networks," Christopher J. Shallue & Daniel J. Eisenstein (2022)</p>
Dataset1 for "General framework for E(3)-equivariant neural network representation of density functional theory Hamiltonian"
<p>Supporting data for the paper "General framework for E(3)-equivariant neural network representation of density functional theory Hamiltonian". </p> <p>Contains atomic structures and Hamiltonian matrices of monolayer graphene, monolayer MoS<sub>2</sub>, bilayer graphene and bilayer bismuthene.</p> <p>Detailed descriptions about the format of data and instructions on how to reproduce the results in the paper can be found in README.md.</p>
Dataset2 for "General framework for E(3)-equivariant neural network representation of density functional theory Hamiltonian"
<p>Supporting data for the paper "General framework for E(3)-equivariant neural network representation of density functional theory Hamiltonian". </p> <p>Contains atomic structures and Hamiltonian matrices of bilayer bismuth selenide. </p> <p>Detailed descriptions about the format of data and instructions on how to reproduce the results in the paper can be found in README.md.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.