Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

558 results for “Training Data”

Learn how ShareScore rates datasets ↗
zenodo36/100

Terms4FAIRskills - an overview for the Research Data Alliance 'Birds of a Feather' session, Skills and training curriculums to support FAIR for Research Software

<p>Short talk (7m) on the terms4FAIRskills initiative for the Research Data Alliance &#39;Birds of a Feather&#39; session, Skills and training curriculums to support FAIR for Research Software, at RDA plenary 18 on 3 Nov 2021.</p> <p>Please see https://www.rd-alliance.org/skills-and-training-curriculums-support-fair-research-software for more information on the BoF session. Please see https://terms4fairskills.github.io/ for more information on the terms4FAIRskills initiative.</p>

opencc-by-sa-4.0Nov 2021View details →
zenodo36/100

Training data for benchtop NMR and UV/vis spectroscopy for Artificial Neural Networks

<p>Data set of low-field NMR spectra and UV/vis spectra for the synthesis of mesalazine intermediates, which were used as training or validation data for data processing with artificial neural networks development</p> <p><strong>Low-field NMR spectra for the nitration step:</strong></p> <p>The pure component spectrum of 2ClBA, 3N-2ClBA, and 5N-2ClBA are marked as NMR_pure_spectrum. The concentration levels for 2ClBA, 3N-2ClBA and 5N-2ClBA are in row 1, 2, and 3, respectively.</p> <p>The data sets marked as NMR_ represents low-field NMR-spectra recorded. The reference values for 2ClBA, 3N-2ClBA and 5N-2ClBA are in column 1, 2, and 3, respectively.</p> <p><strong>Datafusion data sets for the hydrolysis and nitration step</strong></p> <p>The NMR data are either recorded or simulated from the pure NMR spectrum of each individual component. The reference values for 2ClBA, 3N-2ClBA, 5N-2ClBA, 3-NSA and 5-NSA are either assigned with UHPLC measurements or calculated from the prepared solutions.</p> <p>The NMR spectra are depicted in datafusion_NMR_training. The reference values for 2ClBA, 3N-2ClBA and 5N-2ClBA are in column 1, 2, and 3, respectively.</p> <p>The UV/vis spectra are depicted in datafusion_UVvis_training. The reference values for 2ClBA, 3N-2ClBA, 5N-2ClBA, 3-NSA and 5-NSA are in column 1, 2, 3, 4, and 5, respectively.</p> <p><strong>Process data</strong></p> <p>The NMR spectra for the stability run and the run with dynamic changes are depicted in process_NMR_. The first column is the time stamp.</p> <p>The UV/vis spectra for the stability run and the run with dynamic changes are depicted in process_UV_. The first column is the time stamp.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Training dataset for "A deep learned nanowire segmentation model using synthetic data augmentation"

<p>This image dataset contains synthetic structure images used for training the deep-learning based nanowire segmentation model presented in our work &quot;A deep learned nanowire segmentation model using synthetic data augmentation&quot; to be published in <em>npj Computational materials. </em>Detailed information can be found in the corresponding article.</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Documenting And Assessing Open Innovation: Co-creation Of An Open Data Model For Surgical Training (Additional materials, tables 2 & 3)

<p>Challenge competitions have recently resurged for promoting open innovation in areas where markets fail to provide incentives, such as the Sustainable Development Goals (SDGs). Challenges call for the general public to contribute novel solutions to a well-defined problem, in exchange for prizes, credentials and the promise of further development of selected solutions. The aim of this paper is to report on the development of an open and collaborative data model to document and evaluate innovations in the context of a challenge competition, while also being compatible with the work of other open source communities to validate and improve them. By reusing open documentation standards and embedding them into a semantic collaborative platform, the model aimed to be flexible enough to respond to the evaluation needs of the project organisers and self-assessment for participants. We expect our experience provides insights on the potential of semantic, collaborative platforms and standards for increasing the impact of innovations towards the SDGs.</p> <p>The developer team defined the goal and scope of the ontology in collaboration with the GSTC organisers. This was done by agreeing on scenarios where the ontology will be used and establishing competency questions that the ontology has to be able to respond to. Table 2 describes the four motivating scenarios, including actors involved, requirements, sequence of actions and main problems identified. Table 3 details the competency questions for each scenario.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Training and test data, plus saved models for the paper "Top-down effects in an early visual cortex inspired hierarchical Variational Autoencoder" submitted to the SVRHM 2022 Workshop @ NeurIPS

<p>Each .pkl&nbsp;file contains a training or test dataset&nbsp;in the form of a Python dictionary (generated with Python 3.8.5) with the following fields:</p><ul><li>'train_images': 640,000 float32 images&nbsp;used&nbsp;for model training. 20px images contain 400 pixel intensities, 40px images contain 1600 pixel intensities each.</li><li>'train_labels': float32 labels for each image in&nbsp;'train_images'. All natural images are&nbsp;labeled&nbsp;with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0,&nbsp;according to their texture family.</li><li>'test_images': 64,000 float32 images&nbsp;used&nbsp;for model testing.&nbsp;20px images contain 400 pixel intensities, 40px images contain 1600 pixel intensities each.</li><li>'test_labels': float32 labels for each image in&nbsp;'test_images'. All natural images are&nbsp;labeled&nbsp;with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0,&nbsp;according to their texture family.</li></ul><p>Each .zip file contains a saved model.&nbsp;Details on these are coming soon.</p><p>For more details, see the paper&nbsp;"Top-down effects in an early visual cortex inspired hierarchical Variational Autoencoder" published at the SVRHM 2022 Workshop @ NeurIPS&nbsp;(<a href="https://openreview.net/forum?id=8dfboOQfYt3">link</a>).</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Training and validation data for artificial neural networks using three-dimensional partial convolutions to fill gaps in satellite image time series

<p>This dataset contains training and validation data for artificial neural networks using three-dimensional partial convolutions to fill gaps in satellite image time series. The data have been derived from Sentinel-5P total column carbon monoxide observations, using the offline processing stream.</p> <p><strong>Preprocessing</strong></p> <p>The following operations have been applied on the original S5P imagery:</p> <ol> <li>Images have been resampled to 0.1 by 0.1 degree spatial resolution</li> <li>Pixels with quality assessment value less than or equal to 0.5 have been set to NA</li> <li>Images have been aggregated by day of observation</li> <li>Images have been cropped to -60 to 60 degrees latitude</li> <li>Images have been devided into spatiotemporal blocks of size 128 x 128 pixels and 16 days</li> </ol> <p>Imagery has been recorded between 2021-01-01 and 2021-11-25. Notice that both the training and the validation blocks have been randomly sampled from all available blocks.</p> <p><br> <strong>Data Format and Naming Conventions</strong></p> <p>Input and output data blocks are stored as GeoTIFF files, where bands represent time. Notice the following file naming conventions:</p> <ul> <li>Files starting with <em>X</em>&nbsp;represent input measurements for training, where artificial gaps have been added.</li> <li>Files starting with <em>Y</em>&nbsp;represent true measurements without artificially added gaps (but still containing gaps in many cases).</li> <li>Binary masks of input data where all pixels with valid measurements are 1 and others 0 are stored in files whose name starts with <em>MASK</em></li> <li>Files starting with <em>VALMASK</em>&nbsp;contain a binary mask where only pixels that are available in Y but not in X are 1. The latter is used for validation on artificially removed pixels only.</li> </ul> <p>Numbers in filenames encode spatial and temporal block indexes.</p> <p>In addition, the dataset contains prediction of the validation blocks from different models in the `predictions` directory. The subfolders contain output from different models:</p> <ul> <li>mean&nbsp;refers to simple block-wise mean predictions.</li> <li>timeseries&nbsp;refers to simple linear time series interpolation.</li> <li>gapfill&nbsp;refers to the method proposed in [1].</li> <li>stmra&nbsp;refers to the method proposed in [2].</li> <li>STpconv&nbsp;refers to predictions passed on an artificial neural netowork with three-dimensional partial convolutions.</li> </ul> <p><strong>References</strong></p> <p>[1] Gerber, F., de Jong, R., Schaepman, M. E., Schaepman-Strub, G., &amp; Furrer, R. (2018). Predicting missing values in spatio-temporal remote sensing data. IEEE Transactions on Geoscience and Remote Sensing, 56(5), 2841-2853.</p> <p>[2] Appel, M., &amp; Pebesma, E. (2020). Spatiotemporal multi-resolution approximations for analyzing global environmental data. Spatial Statistics, 38, 100465.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Training Data for Decision Tree Practical in reading-ml-chemistry repo

<p>This is a dataset that can be used to train a decision tree to predict the band gap of a material. The data is associated with a notebook for running the practical and it can be found at https://github.com/keeeto/reading-ml-chemistry.</p> <p>&nbsp;</p> <p>The data originally comes from the Materials Project.</p> <p>&nbsp;</p> <p>A new muon spectroscopy dataset is added. It is from - Machine learning approach to muon spectroscopy analysis - https://iopscience.iop.org/article/10.1088/1361-648X/abe39e/meta</p>

opencc-by-4.0Jan 2021View details →
dryad36/100

Ground truth data used to train the synapse classifier used in Lillvis et al., 2022 for ExLLSM circuit reconstruction

<p class="MsoNormal">Brain function is mediated by the physiological coordination of a vast, intricately connected network of molecular and cellular components. The physiological properties of network components can be quantified with high throughput; the ability to assess many animals per study has been key to relating physiological properties to behavior. Conversely, detailed anatomical properties (e.g., the synaptic connectivity of molecularly-defined cell types across an entire circuit) are presently quantifiable only with low throughput; thus we know very little about how network structure, and structural variation, influences behavior. For neuroanatomical reconstruction there is a methodological gulf between electron-microscopic (EM) methods, which yield dense connectomes (but at great expense and low throughput) and light-microscopic methods, which provide molecular and cell-type specificity with high throughput (but without synaptic resolution). We developed a high-throughput analysis pipeline and imaging protocol using tissue expansion and light sheet microscopy (ExLLSM) to rapidly reconstruct selected circuits across many animals with single-synapse resolution and molecular contrast. Using <em>Drosophila </em>to validate this approach, we demonstrate that it yields synaptic counts similar to those obtained by EM, enables synaptic connectivity to be compared across sex and experience, and can be used to correlate structural connectivity, functional connectivity, and behavior. This approach fills a critical methodological gap in studying variability in the structure and function of neural circuits across individuals within and between species.</p> <p class="MsoNormal">Here, we share the data used to train the synapse classifier that was utilized in the analysis pipeline. All additional software, code, and usage examples to train and run the classifier can be found at Github: <a href="https://github.com/JaneliaSciComp/exllsm-circuit-reconstruction">https://github.com/JaneliaSciComp/exllsm-circuit-reconstruction</a></p>

opencc-zeroJul 2022View details →
zenodo36/100

GlottisNetV2 - Time variant training and testing data

<p>Here we provide time variant training and testing data for the GlottisNetV2 study. In particular, this dataset relies on the benchmark for automatic glottis segmentation (BAGLS dataset). This three-dimensional data (time, y, x) is used for experiments involving deep neural networks capable of processing time-variant data.&nbsp;</p> <p>Each folder contains videos each with the following data:</p> <ul> <li>Endoscopic video as mp4 (*.mp4)</li> <li>Glottis segmentation as mask-file (hdf5 container, *.mask)</li> <li>Glottis segmentation as mp4 file (*_mask.mp4)</li> <li>Metadata as JSON file (*.meta)</li> <li>Glottal midline annotation as JSON file (*.points)</li> </ul>

opencc-by-nc-sa-4.0Jul 2022View details →
zenodo36/100

The dataset for an article - An Evaluation of 3D-Printed Materials' Structural Properties Using Active Infrared Thermography and Deep Neural Networks Trained on the Numerical Data

<p>Dataset used in the research presented in the article:</p> <p>Szymanik, Barbara. 2022. &quot;An Evaluation of 3D-Printed Materials&rsquo; Structural Properties Using Active Infrared Thermography and Deep Neural Networks Trained on the Numerical Data&quot;&nbsp;<em>Materials</em>&nbsp;15, no. 10: 3727. https://doi.org/10.3390/ma15103727</p> <p>The database in the .mat (matlab) format contains arrays of double type related to: A - original thermograms obtained for the plate made with the 3D printing technique Ar - thermograms with ROI included FITorg - approximation of original thermograms ImDiff, ImInt, ImProp - data obtained after subtracting the approximation.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Automated systematic evaluation of cryo-EM specimens with SmartScope - Training data for hole detector

<p>Training data for hole detector<br> =======================</p> <p>This dataset includes 36 images.<br> Circles are annotated in COCO format.</p> <p>The following pre-processing was applied to each image:<br> * Auto-orientation of pixel data (with EXIF-orientation stripping)</p> <p>The following augmentation was applied to create 1 versions of each source image:<br> * 50% probability of horizontal flip</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Automated systematic evaluation of cryo-EM specimens with SmartScope - Training data for square detector

<p>Training data for square detector<br> =========================</p> <p>This dataset includes 26 images.<br> Squares are annotated in COCO format.</p> <p>The following pre-processing was applied to each image:<br> * Auto-orientation of pixel data (with EXIF-orientation stripping)</p> <p>The following augmentation was applied to create 1 versions of each source image:<br> * 50% probability of horizontal flip<br> * Equal probability of one of the following 90-degree rotations: none, clockwise, counter-clockwise</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Biotechnology data analysis training with Jupyter Notebooks

<p>Biotechnology has experienced innovations in analytics and data processing. As the volume of data and its complexity grows, new computational procedures for extracting information are developed. However, the rate of change outpaces the adaptation of biotechnology curricula, necessitating new teaching methodologies to equip biotechnologists with data analysis abilities. To simulate experimental data, we created a virtual organism simulator (<em>silvio</em>) by combining diverse cellular and sub-cellular microbial models. With the <em>silvio </em>Python package, we constructed a computer-based instructional workflow to teach growth curve data analysis, promoter sequence design, and expression rate measurement. The instructional workflow is a Jupyter Notebook with background explanations and Python-based experiment simulations combined. The data analysis is either conducted within the Notebook in Python or externally with Excel. This instructional workflow was separately implemented in two distance courses for Master&#39;s students in biology and biotechnology with assessment of the pedagogic efficiency. The concept of using virtual organism simulations that generate coherent results across different experiments can be used to construct consistent and motivating case studies for biotechnological data literacy.</p> <p>Here, the supplementary material is provided.</p> <table> <tbody> <tr> <td>2207_BLS-RecExpSim.mbz</td> <td>Moodle backup file for import as new moodle function.</td> </tr> <tr> <td>BLS_RecExpSim_PerformanceEvaluation Rubric.docx</td> <td>Expected learning outcomes with associated performance levels.</td> </tr> <tr> <td>BLS_SurveryQuestions.docx</td> <td>Survey questions to evaluate the educational approach.</td> </tr> <tr> <td>RecExpSim.html</td> <td>Html-Export of the Jupyter Notebook to teach biotechnology data analysis. This only serves as visual impression of the course because the dynamic Python-evaluations are not functioning.</td> </tr> <tr> <td>RecExpSim_Lecture.pdf</td> <td>Static pdf of preparatory lecture to cover the theoretical aspects in the simulations and to get student on comparable level.</td> </tr> <tr> <td>RecExpSim_Lecture.pptx</td> <td>Adjustable pptx of preparatory lecture to cover the theoretical aspects in the simulations and to get student on comparable level.</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Edited DHS data for R training

<p>Edited DHS data for R training</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Training Data for "DeepCLEM: automated registration for correlative light and electron microscopy using deep learning"

<p><strong>This folder contains the training dataset used for the paper</strong></p> <p>&quot;DeepCLEM: automated registration for correlative light and electron microscopy using deep learning&quot;</p> <p><em>Rick Seifert, Sebastian M. Markert, Sebastian Britz, Veronika Perschin, Christoph Erbacher, Christian Stigloher and Philip Kollmannsberger</em></p> <p>F1000Research 9:1275 (2020), https://f1000research.com/articles/9-1275</p> <p>------------------------------------------------------------</p> <p>These are 117+4 manually aligned CLEM images of C.elegans acquired by Sebastian M. Markert, Sebastian Britz and Rick Seifert in the Electron Microscopy Facility of the Biocenter of University of Wuerzburg, Germany. For details and experimental protocols, please see the paper linked above.</p> <p>Contents:</p> <ul> <li>&quot;fluo_training&quot;: Fluorescence microscopic channel of the 117 training images&nbsp;</li> <li>&quot;sem_training&quot;: Scanning electron microscopic channel of the 117 training images</li> <li>&quot;fluo_validation&quot;: Fluorescence microscopic channel of the 4 validation images&nbsp;</li> <li>&quot;sem_validation&quot;: Scanning electron microscopic channel of the 4 validation images</li> </ul> <p>License: CC-BY 4.0</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Deep Deep Learning With BART (Trained Weights and Example Data)

<p>This repository contains data required to reproduce the figures of the manuscript Deep, Deep Learning with BART. The corresponding scripts can be found at&nbsp;https://github.com/mrirecon/deep-deep-learning-with-bart.</p> <p>The example data in this repository are based on the data published with the manuscripts of the Variational Network [1] and MoDL [2].</p> <p>[1]: Hammernik K, Klatzer T, Kobler E, Recht MP, Sodickson DK, Pock T, Knoll F.<br> Learning a variational network for reconstruction of accelerated MRI data.<br> Magn Reson Med 2018; 79:3055-3071.</p> <p>[2]:&nbsp;Aggarwal HK, Mani MP, Jacob M.<br> MoDL: Model-Based Deep Learning Architecture for Inverse Problems.<br> IEEE Trans Med Imaging 2019; 38:394--405.</p> <p>&nbsp;</p>

opencc-by-nc-4.0Apr 2022View details →
zenodo36/100

Synthetic data for R training

<p>Synthetic data for R training</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Training Data for 'somatic variants discovery'

<p>The data provided here are part of a Galaxy Training Network tutorial created by&nbsp;Bj&ouml;rn Gr&uuml;ning&nbsp;to&nbsp;detect hCNVS in defined WES chromosomal&nbsp;regions.</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

NASA SPoRT Basin Average Training Data

<p>The included data files contain basin average SPoRT-LIS relative soil moisture [Total column (0-2 m depth) and four model layers (0-10, 10-40, 40-100, and 100-200 cm depth)] and&nbsp;MRMS QPE for each river basin. These files were used to train and tune the developed basin specific LSTM models. The number in the file name&nbsp;indicates&nbsp;the corresponding USGS site number.&nbsp;&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Training data and models for microphysics emulation

<p>Training data and models for microphysics emulation</p> <p>The training data is a subset of the full dataset described in the NeurIPS submission. Roughly speaking, 30 day runs with FV3GFS, with zhao carr microphysics. To keep the data reasonable in size, 1000 random netCDFs are sampled from the over 7000 files in the full training dataset. 200 test files are sampled.</p> <p>Also contains the trained ML models at models/</p> <p>Data behind the plots and tables is at plot-data/.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record